Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 222 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Neural Network Models for Contextual Regression
Seksan Kiatsupaibul, Pakawan Chansiripas · 26 mars 2026
We propose a neural network model for contextual regression in which the regression model depends on contextual features that determine the active submodel and an algorithm to fit the model. The proposed simple contextual neural network (SCtxtNN) separates context identification from context-specifi…
- Minimal Sufficient Representations for Self-interpretable Deep Neural Networks
Zhiyao Tan, Liu Li, Huazhen Lin · 26 mars 2026
Deep neural networks (DNNs) achieve remarkable predictive performance but remain difficult to interpret, largely due to overparameterization that obscures the minimal structure required for interpretation. Here we introduce DeepIn, a self-interpretable neural network framework that adaptively identi…
- No Single Metric Tells the Whole Story: A Multi-Dimensional Evaluation Framework for Uncertainty Attributions
Emily Schiller, Teodor Chiaburu, Marco Zullich, Luca Longo · 26 mars 2026
Research on explainable AI (XAI) has frequently focused on explaining model predictions. More recently, methods have been proposed to explain prediction uncertainty by attributing it to input features (uncertainty attributions). However, the evaluation of these methods remains inconsistent as studie…
- BXRL: Behavior-Explainable Reinforcement Learning
Ram Rachum, Yotam Amitai, Yonatan Nakar, Reuth Mirsky, Cameron Allen · 26 mars 2026
A major challenge of Reinforcement Learning is that agents often learn undesired behaviors that seem to defy the reward structure they were given. Explainable Reinforcement Learning (XRL) methods can answer queries such as "explain this specific action", "explain this specific trajectory", and "expl…
- Leakage and Interpretability in Concept-Based Models
Enrico Parisini, Tapabrata Chakraborti, Chris Harbron, Ben D. MacArthur, Christopher R. S. Banerji · 25 mars 2026
Concept-based Models aim to improve interpretability by predicting high-level intermediate concepts, representing a promising approach for deployment in high-risk scenarios. However, they are known to suffer from information leakage, whereby models exploit unintended information encoded within the l…
- Toward Faithful Segmentation Attribution via Benchmarking and Dual-Evidence Fusion
Abu Noman Md Sakib, OFM Riaz Rahman Aranya, Kevin Desai, Zijie Zhang · 25 mars 2026
Attribution maps for semantic segmentation are almost always judged by visual plausibility. Yet looking convincing does not guarantee that the highlighted pixels actually drive the model's prediction, nor that attribution credit stays within the target region. These questions require a dedicated eva…
- Biased Error Attribution in Multi-Agent Human-AI Systems Under Delayed Feedback
Teerthaa Parakh, Karen M. Feigh · 25 mars 2026
Human decision-making is strongly influenced by cognitive biases, particularly under conditions of uncertainty and risk. While prior work has examined bias in single-step decisions with immediate outcomes and in human interaction with a single autonomous agent, comparatively little attention has bee…
- AI Mental Models: Learned Intuition and Deliberation in a Bounded Neural Architecture
Laurence Anthony · 25 mars 2026
This paper asks whether a bounded neural architecture can exhibit a meaningful division of labor between intuition and deliberation on a classic 64-item syllogistic reasoning benchmark. More broadly, the benchmark is relevant to ongoing debates about world models and multi-stage reasoning in AI. It …
- Language Models Can Explain Visual Features via Steering
Javier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres, Daniel Hinjos, Anna Arias-Duart, Dario Garcia-Gasulla · 25 mars 2026
Sparse Autoencoders uncover thousands of features in vision models, yet explaining these features without requiring human intervention remains an open challenge. While previous work has proposed generating correlation-based explanations based on top activating input examples, we present a fundamenta…
- YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception
Marios Impraimakis, Daniel Vazquez, Feiyu Zhou · 25 mars 2026
The interpretable object detection capabilities of a novel Kolmogorov-Arnold network framework are examined here. The approach refers to a key limitation in computer vision for autonomous vehicles perception, and beyond. These systems offer limited transparency regarding the reliability of their con…
- First-Mover Bias in Gradient Boosting Explanations: Mechanism, Detection, and Resolution
Drake Caraker, Bryan Arnold, David Rhoads · 25 mars 2026
We isolate and empirically characterize first-mover bias -- a path-dependent concentration of feature importance caused by sequential residual fitting in gradient boosting -- as a specific mechanistic cause of the well-known instability of SHAP-based feature rankings under multicollinearity. When co…
- Reasoning Traces Shape Outputs but Models Won't Say So
Yijie Hao, Lingjie Chen, Ali Emami, Joyce Ho · 24 mars 2026
Can we trust the reasoning traces that large reasoning models (LRMs) produce? We investigate whether these traces faithfully reflect what drives model outputs, and whether models will honestly report their influence. We introduce Thought Injection, a method that injects synthetic reasoning snippets …
- How AI Systems Think About Education: Analyzing Latent Preference Patterns in Large Language Models
Daniel Autenrieth · 24 mars 2026
This paper presents the first systematic measurement of educational alignment in Large Language Models. Using a Delphi-validated instrument comprising 48 items across eight educational-theoretical dimensions, the study reveals that GPT-5.1 exhibits highly coherent preference patterns (99.78% transit…
- Position: Multi-Agent Algorithmic Care Systems Demand Contestability for Trustworthy AI
Truong Thanh Hung Nguyen, H\'el\`ene Fournier, Piper Jackson, Makoto Itoh, Shannon Freeman, Rene Richard, Hung Cao · 24 mars 2026
Multi-agent systems (MAS) are increasingly used in healthcare to support complex decision-making through collaboration among specialized agents. Because these systems act as collective decision-makers, they raise challenges for trust, accountability, and human oversight. Existing approaches to trust…
- Does Mechanistic Interpretability Transfer Across Data Modalities? A Cross-Domain Causal Circuit Analysis of Variational Autoencoders
Dip Roy, Rajiv Misra, Sanjay Kumar Singh, Anisha Roy · 24 mars 2026
Although mechanism-based interpretability has generated an abundance of insight for discriminative network analysis, generative models are less understood -- particularly outside of image-related applications. We investigate how much of the causal circuitry found within image-related variational aut…
- Enhancing LIME using Neural Decision Trees
Mohamed Aymen Bouyahia, Argyris Kalogeratos · 24 mars 2026
Interpreting complex machine learning models is a critical challenge, especially for tabular data where model transparency is paramount. Local Interpretable Model-Agnostic Explanations (LIME) has been a very popular framework for interpretable machine learning, also inspiring many extensions. While …
- Me, Myself, and $\pi$ : Evaluating and Explaining LLM Introspection
Atharv Naphade, Samarth Bhargav, Sean Lim, Mcnair Shah · 24 mars 2026
A hallmark of human intelligence is Introspection-the ability to assess and reason about one's own cognitive processes. Introspection has emerged as a promising but contested capability in large language models (LLMs). However, current evaluations often fail to distinguish genuine meta-cognition fro…
- The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
Shasha Yu, Fiona Carroll, Barry L. Bentley · 24 mars 2026
Large language models (LLMs) are increasingly deployed as agents with access to executable tools, enabling direct interaction with external systems. However, most safety evaluations remain text-centric and assume that compliant language implies safe behavior, an assumption that becomes unreliable on…
- ShapDBM: Exploring Decision Boundary Maps in Shapley Space
Luke Watkin, Daniel Archambault, Alex Telea · 24 mars 2026
Decision Boundary Maps (DBMs) are an effective tool for visualising machine learning classification boundaries. Yet, DBM quality strongly depends on the dimensionality reduction (DR) technique and high dimensional space used for the data points. For complex ML datasets, DR can create many mixed clas…
- On Arbitrary Predictions from Equally Valid Models
Sarah Lockfisch, Kristian Schwethelm, Martin Menten, Rickmer Braren, Daniel Rueckert, Alexander Ziller, Georgios Kaissis · 24 mars 2026
Model multiplicity refers to the existence of multiple machine learning models that describe the data equally well but may produce different predictions on individual samples. In medicine, these models can admit conflicting predictions for the same patient -- a risk that is poorly understood and ins…
- TREX: Trajectory Explanations for Multi-Objective Reinforcement Learning
Dilina Rajapakse, Juan C. Rosero, Ivana Dusparic · 24 mars 2026
Reinforcement Learning (RL) has demonstrated its ability to solve complex decision-making problems in a variety of domains, by optimizing reward signals obtained through interaction with an environment. However, many real-world scenarios involve multiple, potentially conflicting objectives that cann…
- Interpretable Multiple Myeloma Prognosis with Observational Medical Outcomes Partnership Data
Salma Rachidi, Aso Bozorgpanah, Eric Fey, Alexander Jung · 24 mars 2026
Machine learning (ML) promises better clinical decision-making, yet opaque model behavior limits the adoption in healthcare. We propose two novel regularization techniques for ensuring the interpretability of ML models trained on real-world data. In particular, we consider the prediction of five-yea…
- CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
Florent Draye, Abir Harrasse, Vedant Palit, Tung-Yu Wu, Jiarui Liu, Punya Syon Pandey, Roderick Wu, Terry Jingchen Zhang, Zhijing Jin, Bernhard Sch\"olkopf · 24 mars 2026
Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information. Recent approaches based on dictionary learning and transcoders enable representing model computation in terms of sparse, interpretable features and their interactions, giving rise to …
- GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models
Tianyu Bell Pan, Damon L. Woodard · 23 mars 2026
Large language models (LLMs) demonstrate strong performance, but they often lack transparency. We introduce GeoLAN, a training framework that treats token representations as geometric trajectories and applies stickiness conditions inspired by recent developments related to the Kakeya Conjecture. We …
- Pitfalls in Evaluating Interpretability Agents
Tal Haklay, Nikhil Prakash, Sana Pandey, Antonio Torralba, Aaron Mueller, Jacob Andreas, Tamar Rott Shaham, Yonatan Belinkov · 23 mars 2026
Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverage large language models (LLMs) at increasing levels of autonomy, ranging from fixed one-shot workflows to fully autonomo…
