Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 222 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- AttnGen: Attention-Guided Saliency Learning for Interpretable Genomic Sequence Classification
Rayhaneh Shabani Nia, Ali Karkehabadi · 15 mai 2026
Deep neural networks have achieved strong performance in genomic sequence classification; however, relating their predictions to biologically meaningful sequence patterns remains challenging. In this work, we present AttnGen, an attention-guided training framework that embeds interpretability direct…
- Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity
Cristian Hinostroza, Rodrigo Toro Icarte, Christ Devia, Andres Carvallo De Ferari, Eugenio Herrera-Berg, Denis Parra, Jorge F Silva · 15 mai 2026
Large language models (LLMs) have revolutionized natural language processing. Understanding their internal mechanisms is crucial for developing more interpretable and optimized architectures. Mechanistic interpretability has led to the development of various methods for assessing layer relevance, wi…
- The Rate-Distortion-Polysemanticity Tradeoff in SAEs
Tommaso Mencattini, Francesco Montagna, Francesco Locatello · 15 mai 2026
Sparse Autoencoders (SAEs) that can accurately reconstruct their input (minimizing distortion) by making efficient use of few features (minimizing the rate) often fail to learn monosemantic representations (highly interpretable), limiting their usefulness for mechanistic interpretability. In this pa…
- Finding Interpretable Prompt-Specific Circuits in Language Models
Gabriel Franco, Lucas M. Tassis, Azalea Rohr, Mark Crovella · 15 mai 2026
Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A crucial part of finding circuits is understanding why each attention head attends where it does. To this end, we introduce ACC++, an improved circuit-tracing met…
- AIMing for Standardised Explainability Evaluation in GNNs: A Framework and Case Study on Graph Kernel Networks
Magdalena Proszewska, N. Siddharth · 15 mai 2026
Graph Neural Networks (GNNs) have advanced significantly in handling graph-structured data, but a comprehensive framework for evaluating explainability remains lacking. Existing evaluation frameworks primarily involve post-hoc explanations, and operate in the setting where multiple methods generate …
- Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
Adarsh Kumarappan, Ananya Mujoo · 15 mai 2026
LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced sycophancy. We test this attribution across four model families and find it largely wrong: pretrained base models exhibit…
- Architecture-Aware Explanation Auditing for Industrial Visual Inspection
Sibo Jia, Zihang Zhao, Kunrong Li · 15 mai 2026
Industrial visual inspection systems increasingly rely on deep classifiers whose heatmap explanations may appear visually plausible while failing to identify the image regions that actually drive model decisions. This paper operationalizes an architecture-aware explanation audit protocol grounded in…
- Woodelf++: A Fast and Unified Partial Dependence Plot Algorithm for Decision Tree Ensembles
Ron Wettenstein, Alexander Nadel, Udi Boker · 15 mai 2026
Partial Dependence Plots (PDPs) visualize how changes in a single feature affect the average model prediction. They are widely used in practice to interpret decision tree ensembles and other machine learning models. Joint-PDPs extend this idea to pairs of features, revealing their combined effect. P…
- Enhanced and Efficient Reasoning in Large Learning Models
Leslie G. Valiant · 15 mai 2026
In current Large Language Models we can trust the production of smoothly flowing prose on the basis of the principles of machine learning. However, there is no comparably principled basis to justify trust in the content of the text produced. It appears to be conventional wisdom that addressing this …
- RoSHAP: A Distributional Framework and Robust Metric for Stable Feature Attribution
Lanxin Xiang, Liang Shi, Youhui Ye, Boyu Jiang, Dawei Zhou, Feng Guo · 15 mai 2026
Feature attribution analysis is critical for interpreting machine learning models and supporting reliable data-driven decisions. However, feature attribution measures often exhibit stochastic variation: different train--test splits, random seeds, or model-fitting procedures can produce substantially…
- Exemplar Partitioning for Mechanistic Interpretability
Jessica Rumbelow · 15 mai 2026
We introduce Exemplar Partitioning (EP), an unsupervised method for constructing interpretable feature dictionaries from large language model activations with $\sim 10^{3}\times$ fewer tokens than comparable sparse autoencoders (SAEs). An EP dictionary is a Voronoi partition of activation space, bui…
- Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models
Senne Deproost, Denis Steckelmacher, Ann Now\'e · 15 mai 2026
Despite many successful attempts at explaining Deep Reinforcement Learning policies using distillation, it remains difficult to balance the performance-interpretability trade-off and select a fitting surrogate model. In addition to this, traditional distillation only minimizes the distance between t…
- From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
Jinxian Qu, Qingqing Gu, Teng Chen, Luo Ji · 15 mai 2026
Wide applications of LLM-based agents require strong alignment with human social values. However, current works still exhibit deficiencies in self-cognition and dilemma decision, as well as self-emotions. To remedy this, we propose a novel value-based framework that employs GraphRAG to convert princ…
- Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems
Jos\'e Manuel de la Chica Rodr\'iguez, Carlos Mart\'i-Gonz\'alez · 15 mai 2026
Large language models in regulated financial workflows are governed by natural-language policies that the same model interprets, creating a principal--agent failure: outputs can appear compliant without being compliant. Existing evaluation measures task accuracy but not whether governance constrains…
- LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling
Qi Cao, Yufan Wang, Peijia Qin, Shuhao Zhang, Pengtao Xie · 15 mai 2026
Large language models (LLMs) often expose useful signals of self-monitoring: before solving a problem, they can estimate whether they are likely to succeed, and after solving it, they can judge whether their answer is likely to be correct. However, these signals are typically measured or elicited in…
- Towards Fine-Grained and Verifiable Concept Bottleneck Models
Yingying Fang, Haijie Xu, Shuang Wu, Mariathasan Anish, Guang Yang · 15 mai 2026
Concept Bottleneck Models (CBMs) offer interpretable alternatives to black-box predictors by introducing human-relatable concepts before the final output. However, existing CBMs struggle to verify whether predicted concepts correspond to the correct visual evidence, limiting their reliability. We pr…
- Wahkon: A Statistically Principled Deep RKHS Superposition Network
Yongkai Chen, Wenxuan Zhong, Ping Ma · 15 mai 2026
Deep learning excels at prediction but often lacks finite-sample guarantees and calibrated uncertainty; RKHS (Reproducing Kernel Hilbert Space)-based methods provide those guarantees but struggle to adapt in high dimensions. We propose Wahkon, a deep RKHS superposition network that unifies Kolmogoro…
- Multi-Dimensional Model Integrity and Responsibility Assessment Index and Scoring Framework
Phuc Truong Loc Nguyen, Thanh Hung Do, Truong Thanh Hung Nguyen, Hung Cao · 15 mai 2026
Artificial intelligence in high-stakes tabular domains cannot be evaluated by predictive performance alone, yet current practice still assesses explainability, fairness, robustness, privacy, and sustainability mostly in isolation. We propose the Model Integrity and Responsibility Assessment Index (M…
- Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification
One Octadion, Novanto Yudistira, Lailil Muflikhah · 14 mai 2026
Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Prediction (CP) provides distribution-free statistical guarantees, standard methods such as Regularized Adaptive Prediction Sets (RAPS) optimize for average…
- AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
Raj Kiran Gupta Katakam · 14 mai 2026
The Average Gradient Outer Product (AGOP) governs feature learning in neural networks: the Neural Feature Ansatz states that weight Gram matrices at each layer align with the corresponding AGOP matrices computed over the training distribution. We ask a complementary question: can this same quantity …
- Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
Adarsh Kumarappan, Ananya Mujoo · 14 mai 2026
LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced sycophancy. We test this attribution across four model families and find it largely wrong: pretrained base models exhibit…
- MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
H. Moore, S. Qi, D. Milojicic, C. Bash, S. Pasricha · 14 mai 2026
Large Language Models (LLMs) have become increasingly prevalent in cloud-based platforms, propelled by the introduction of AI-based consumer and enterprise services. LLM inference requests in particular account for up to 90% of total LLM lifecycle energy use, dwarfing training energy costs. The risi…
- Where Does Reasoning Break? Step-Level Hallucination Detection via Hidden-State Transport Geometry
Tyler Alvarez, Ali Baheri · 14 mai 2026
Large language models hallucinate during multi-step reasoning, but most existing detectors operate at the trace level: they assign one confidence score to a full output, fail to localize the first error, and often require multiple sampled completions. We frame hallucination instead as a property of …
- Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
Jordan F. McCann · 14 mai 2026
Sparse autoencoders (SAEs) are now standard tools for decomposing language model activations into interpretable features, and automated interpretability pipelines routinely assign each feature a short natural-language explanation. Existing critiques of this practice focus on polysemanticity -- one f…
- No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
Ying Li, Hongbo Wen, Yanju Chen, Hanzhi Liu, Yuan Tian, Yu Feng · 14 mai 2026
LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, but because the skill it invoked broke its own declared safety rules. We call these specification violations: benign inputs cause a skill to breach the…
