Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 226 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Preserving Task-Relevant Information Under Linear Concept Removal
Floris Holstege, Shauli Ravfogel, Bram Wouters · 17 novembre 2025
Modern neural networks often encode unwanted concepts alongside task-relevant information, leading to fairness and interpretability concerns. Existing post-hoc approaches can remove undesired concepts but often degrade useful signals. We introduce SPLINCE-Simultaneous Projection for LINear concept r…
- ICX360: In-Context eXplainability 360 Toolkit
Dennis Wei, Ronny Luss, Xiaomeng Hu, Lucas Monteiro Paes, Pin-Yu Chen, Karthikeyan Natesan Ramamurthy, Erik Miehling, Inge Vejsbjerg, Hendrik Strobelt · 17 novembre 2025
Large Language Models (LLMs) have become ubiquitous in everyday life and are entering higher-stakes applications ranging from summarizing meeting transcripts to answering doctors' questions. As was the case with earlier predictive models, it is crucial that we develop tools for explaining the output…
- Higher-order Neural Additive Models: An Interpretable Machine Learning Model with Feature Interactions
Minkyu Kim, Hyun-Soo Choi, Jinho Kim · 17 novembre 2025
Neural Additive Models (NAMs) have recently demonstrated promising predictive performance while maintaining interpretability. However, their capacity is limited to capturing only first-order feature interactions, which restricts their effectiveness on real-world datasets. To address this limitation,…
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
Thomas Decker, Volker Tresp, Florian Buettner · 14 novembre 2025
Perturbation-based explanations are widely utilized to enhance the transparency of machine-learning models in practice. However, their reliability is often compromised by the unknown model behavior under the specific perturbations used. This paper investigates the relationship between uncertainty ca…
- Belief Net: A Filter-Based Framework for Learning Hidden Markov Models from Observations
Reginald Zhiyan Chen, Heng-Sheng Chang, Prashant G. Mehta · 14 novembre 2025
Hidden Markov Models (HMMs) are fundamental for modeling sequential data, yet learning their parameters from observations remains challenging. Classical methods like the Baum-Welch (EM) algorithm are computationally intensive and prone to local optima, while modern spectral algorithms offer provable…
- DenoGrad: Deep Gradient Denoising Framework for Enhancing the Performance of Interpretable AI Models
J. Javier Alonso-Ramos, Ignacio Aguilera-Martos, Andr\'es Herrera-Poyatos, Francisco Herrera · 14 novembre 2025
The performance of Machine Learning (ML) models, particularly those operating within the Interpretable Artificial Intelligence (Interpretable AI) framework, is significantly affected by the presence of noise in both training and production data. Denoising has therefore become a critical preprocessin…
- Privacy-Preserving Explainable AIoT Application via SHAP Entropy Regularization
Dilli Prasad Sharma, Xiaowei Sun, Liang Xue, Xiaodong Lin, Pulei Xiong · 14 novembre 2025
The widespread integration of Artificial Intelligence of Things (AIoT) in smart home environments has amplified the demand for transparent and interpretable machine learning models. To foster user trust and comply with emerging regulatory frameworks, the Explainable AI (XAI) methods, particularly po…
- Efficiently Transforming Neural Networks into Decision Trees: A Path to Ground Truth Explanations with RENTT
Helena Monke, Benjamin Fresz, Marco Bernreuther, Yilin Chen, Marco F. Huber · 13 novembre 2025
Although neural networks are a powerful tool, their widespread use is hindered by the opacity of their decisions and their black-box nature, which result in a lack of trustworthiness. To alleviate this problem, methods in the field of explainable Artificial Intelligence try to unveil how such automa…
- From Decision Trees to Boolean Logic: A Fast and Unified SHAP Algorithm
Alexander Nadel, Ron Wettenstein · 13 novembre 2025
SHapley Additive exPlanations (SHAP) is a key tool for interpreting decision tree ensembles by assigning contribution values to features. It is widely used in finance, advertising, medicine, and other domains. Two main approaches to SHAP calculation exist: Path-Dependent SHAP, which leverages the tr…
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
Laura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer, Anna Hedstr\"om, Marina M. -C. H\"ohne, Oliver Eberle · 13 novembre 2025
Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face…
- Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
Ege Erdogan, Ana Lucic · 13 novembre 2025
Sparse autoencoders (SAEs) have proven useful in disentangling the opaque activations of neural networks, primarily large language models, into sets of interpretable features. However, adapting them to domains beyond language, such as scientific data with group symmetries, introduces challenges that…
- Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
Marios Kokkodis, Richard Demsyn-Jones, Vijay Raghavan · 13 novembre 2025
Are traditional classification approaches irrelevant in this era of AI hype? We show that there are multiclass classification problems where predictive models holistically outperform LLM prompt-based frameworks. Given text and images from home-service project descriptions provided by Thumbtack custo…
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel · 13 novembre 2025
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level algorithm if there exists a function which allows us to map between them. Notably, most interpretabi…
- Decomposition of Small Transformer Models
Casper L. Christensen, Logan Riggs · 13 novembre 2025
Recent work in mechanistic interpretability has shown that decomposing models in parameter space may yield clean handles for analysis and intervention. Previous methods have demonstrated successful applications on a wide range of toy models, but the gap to "real models" has not yet been bridged. In …
- Benevolent Dictators? On LLM Agent Behavior in Dictator Games
Andreas Einwiller, Kanishka Ghosh Dastidar, Artur Romazanov, Annette Hautli-Janisz, Michael Granitzer, Florian Lemmerich · 13 novembre 2025
In behavioral sciences, experiments such as the ultimatum game are conducted to assess preferences for fairness or self-interest of study participants. In the dictator game, a simplified version of the ultimatum game where only one of two players makes a single decision, the dictator unilaterally de…
- Distribution-Based Feature Attribution for Explaining the Predictions of Any Classifier
Xinpeng Li, Kai Ming Ting · 13 novembre 2025
The proliferation of complex, black-box AI models has intensified the need for techniques that can explain their decisions. Feature attribution methods have become a popular solution for providing post-hoc explanations, yet the field has historically lacked a formal problem definition. This paper ad…
- DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
Xueliang Zhao, Wei Wu, Jian Guan, Qintong Li, Lingpeng Kong · 12 novembre 2025
In modern sequential decision-making systems, the construction of an optimal candidate action space is critical to efficient inference. However, existing approaches either rely on manually defined action spaces that lack scalability or utilize unstructured spaces that render exhaustive search comput…
- Rethinking Explanation Evaluation under the Retraining Scheme
Yi Cai, Thibaud Ardoin, Mayank Gulati, Gerhard Wunder · 12 novembre 2025
Feature attribution has gained prominence as a tool for explaining model decisions, yet evaluating explanation quality remains challenging due to the absence of ground-truth explanations. To circumvent this, explanation-guided input manipulation has emerged as an indirect evaluation strategy, measur…
- Data Descriptions from Large Language Models with Influence Estimation
Chaeri Kim, Jaeyeon Bae, Taehwan Kim · 12 novembre 2025
Deep learning models have been successful in many areas but understanding their behaviors still remains a black-box. Most prior explainable AI (XAI) approaches have focused on interpreting and explaining how models make predictions. In contrast, we would like to understand how data can be explained …
- SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto, Fabio Roli, Battista Biggio · 12 novembre 2025
Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model's latent space; e.g., computed …
- Training Language Models to Explain Their Own Computations
Belinda Z. Li, Zifan Carl Guo, Vincent Huang, Jacob Steinhardt, Jacob Andreas · 12 novembre 2025
Can language models (LMs) learn to faithfully describe their internal computations? Are they better able to describe themselves than other models? We study the extent to which LMs' privileged access to their own internals can be leveraged to produce new techniques for explaining their behavior. Usin…
- From Confusion to Clarity: ProtoScore - A Framework for Evaluating Prototype-Based XAI
Helena Monke, Benjamin Sae-Chew, Benjamin Fresz, Marco F. Huber · 12 novembre 2025
The complexity and opacity of neural networks (NNs) pose significant challenges, particularly in high-stakes fields such as healthcare, finance, and law, where understanding decision-making processes is crucial. To address these issues, the field of explainable artificial intelligence (XAI) has deve…
- Hierarchical Deep Counterfactual Regret Minimization
Jiayu Chen, Zhekai Wang, Vaneet Aggarwal · 12 novembre 2025
Imperfect Information Games (IIGs) offer robust models for scenarios where decision-makers face uncertainty or lack complete information. Counterfactual Regret Minimization (CFR) has been one of the most successful family of algorithms for tackling IIGs. The integration of skill-based strategy learn…
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
Jing Huang, Junyi Tao, Thomas Icard, Diyi Yang, Christopher Potts · 12 novembre 2025
Interpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will behave on out-of-distribution examples? In this work, we provide a positive answer to this question. Through a diverse …
- Interpretable Reward Model via Sparse Autoencoder
Shuyi Zhang, Wei Shi, Sihang Li, Jiayi Liao, Tao Liang, Hengxing Cai, Xiang Wang · 12 novembre 2025
Large language models (LLMs) have been widely deployed across numerous fields. Reinforcement Learning from Human Feedback (RLHF) leverages reward models (RMs) as proxies for human preferences to align LLM behaviors with human values, making the accuracy, reliability, and interpretability of RMs crit…
