Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 219 papiers indexés
Volume mensuel — 12 derniers mois
Derniers papiers
- Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls
Jiancu Chen, Shuyin Xia, Guan Wang, Degang Chen, Fan Chen · 24 juillet 2026
Instance-level explanations aim to reveal the rationale behind a model's decisions for a specific graph. Previous methods explain graph neural networks (GNNs) by selecting important edges to induce subgraphs, where edge importance is assessed by perturbing each edge and observing changes in the mode…
- pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development
Chen Zhu, Xiaolu Wang, Weilong Zhang · 24 juillet 2026
In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists. This creates a distinctive reliability problem for multi-agent systems: how should generation, critique, coordination, and hu…
- Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
Moshiur Rahman Tonmoy, Dunren Che, Haitham Y. Adarbah, Afzel Noore · 24 juillet 2026
Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in medical imaging, their trustworthiness is often limited by the quality and granularity of available supervision. In particular, predicted concept activations can…
- Bayesian uncertainty estimation improves clinical decision making in medical AI agents
Frederik Hauke, Patrick Wienholt, Christiane Kuhl, Dyke Ferber, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn · 24 juillet 2026
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to a multi-task chest-radiograph classifier (eight thoracic findings, 137,593 training images), provid…
- When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
Todd Zhou · 24 juillet 2026
Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study this pass@k inversion: after training, the policy may solve fewer distinct problems than its base model at large $k$. The failure concentrates on boundar…
- Gradient Concentration, Not Weight Saliency, Explains Representation-Level Class Unlearning
Billel Habbati, Alessio Merlo, Luca Verderame, Meriem Guerar · 24 juillet 2026
Machine unlearning aims to remove the influence of specific training data while preserving model utility. Many state-of-the-art approaches pursue this goal by restricting the forgetting update to a subset of parameters selected through gradient-based saliency. Although such methods are widely adopte…
- AREX: Towards a Recursively Self-Improving Agent for Deep Research
Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu · 24 juillet 2026
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do mo…
- Explainability Framework for Policy-Aware Autonomous Agents
Heather Merhout (Miami University), Daniela Inclezan (Miami University) · 24 juillet 2026
In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal. As these systems grow more prevalent in our day-to-day lives, there has been an increased need to add explainability features which can provide an account for …
- DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making
Raffi Khatchadourian · 24 juillet 2026
Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We introduce DFAH-Bench, a replay benchmark that measures observable behavioral instability in financial agent decision-making across three channels -- …
- Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement
Jiawei Zheng, Jiazhen Zhang · 24 juillet 2026
Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality. This assumption is poorly sui…
- Statistical Inference for Generative Model Comparison
Zijun Gao, Yan Sun, Han Su · 24 juillet 2026
Generative models have achieved remarkable success across a range of applications, yet their evaluation still lacks principled uncertainty quantification. In this paper, we develop a method for comparing how close different generative models are to the underlying distribution of test samples. Partic…
- Training Large Language Models for Self-Explanation Faithfulness
Yeoktatt Cheah, Mar\'ia P\'erez-Ortiz, Noah Y. Siegel, Oana-Maria Camburu · 24 juillet 2026
We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects its internal decision-making process. While existing work focuses on evaluating faithfulness or using inference-time prom…
- Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity
Hongnan Ma, Yiwei Shi, Mengyue Yang, Weiru Liu · 24 juillet 2026
Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for maintaining it. However, existing sufficiency-oriented methods can assign high importance to spurious subsequences that suppor…
- Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
Seonglae Cho, Zekun Wu, Kleyton Da Costa, Rishi Kalra, Ilham Wicaksono, Adriano Koshiyama · 24 juillet 2026
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE families remains untested. Single-token features that activate on one vocabulary item provide the diagnostic case where ground truth permits direct compar…
- Scaling Interpretable Transformers with Parity Bottleneck Layers
Andrew Mack, Kraig Yuheng Tou, Mark Henry, Zhengxun Wu, Lauren Greenspan · 24 juillet 2026
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are designed to recover such features post-hoc, but training models that are interpretable by construction has remained imprac…
- Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog
Bartolomeo Bogliolo · 24 juillet 2026
Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approaches address this gap by coupling neural models with external symboli…
- Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
Sagar Dangal, Manoj Shakya · 24 juillet 2026
Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval remains untested. We present six coordinated experiments across GPT-2, LLaMA-3.2-1B/3B, OPT-1.3B, and distilgpt2. We first characterise short-term (5-10…
- The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
Yu Wang · 24 juillet 2026
Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow. We show that under group-normalized RL (GRPO), this recipe does not merely fail -- it destroys the policy. Across Qwen3-1.7B/4…
- White Box Evidence Packages for Policy Audit Reports
Seunghyun Yoo · 24 juillet 2026
As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that question in passage-anchored policy audits, where a report must interpret a given pol…
- Expectation Alignment of Language Models for Real-World User Expectations
Miaomiao Li, Yang Wang, Bin Liang, Shudong Liu, Zhiwei Zhang, Kam-Fai Wong · 24 juillet 2026
Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model heuristics, expert rubrics, or user simulation, fail to capture the diversity an…
- Reliability-Aware LLM Alignment from Inconsistent Human Feedback
Jingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma, Lanyu Shang, Yang Zhang · 24 juillet 2026
Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, its efficacy is often compromised by the inherent inconsistency and subjectivity of human annotations. Existing preference optimization frameworks, such as Direct …
- Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady · 23 juillet 2026
Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores must reflect stable propert…
- Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking
Ralf Raumanns, Theresa Elstner, Louis Ferger-Andrews, Louise M. Carlsen, Martin Potthast, Gerard Schouten, Josien P. W. Pluim, Veronika Cheplygina · 23 juillet 2026
Machine learning courses often use pre-labeled datasets, hiding the subjectivity of human annotation. This creates students with an overly trusting view of AI data and models, undervaluing interpretive diversity. We investigated whether manual data annotation tasks teach students about subjective la…
- The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks
Antonio Di Cecco · 23 juillet 2026
Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. We introduce the quadrilateral loss, a differentiable penalty that treats additivity as a measurable behavior instead: a second-order mixed difference on pairs of…
- Self-Explaining Reinforcement Learning for Mobile Network Resource Allocation
Konrad Nowosadko, Franco Ruggeri, Ahmad Terra · 23 juillet 2026
Deep reinforcement learning (DRL) methods, though powerful, often lack transparency, which limits their adoption in critical domains. We apply Self-Explaining Neural Networks (SENNs) to RL by parametrizing the policy of a PPO agent with a SENN, producing intrinsic local explanations, and propose a m…