Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 219 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
Andreas Kouridakis, Dimitrios Patiniotis Spyropoulos, George Vouros · 7 juillet 2026
The use of Large Language Models (LLMs) across diverse areas of human activity-ranging from everyday tasks to safety-critical applications-aims to enhance decision-making effectiveness with minimal human feedback. Concurrently, it seeks to align decisions with human expectations, preferences, and ne…
- Individual Parameters in Weight-Sparse Transformers Appear Interpretable
Arnau Marin-Llobet, Stefan Heimersheim · 7 juillet 2026
A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does. Dominant circuit-finding approaches focus on a specific behavior and reverse-engineer the role of components on the associated sub-distribution. However, past work has sh…
- EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Deyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi · 7 juillet 2026
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to …
- Explainable Bayesian deep learning through input-skip Latent Binary Bayesian Neural Networks
Eirik H{\o}yheim, Lars Skaaret-Lund, Solve S{\ae}b{\o}, Aliaksandr Hubin · 7 juillet 2026
Modeling natural phenomena with artificial neural networks (ANNs) often provides highly accurate predictions. However, ANNs often suffer from over-parameterization, complicating interpretation and raising uncertainty issues. Bayesian neural networks (BNNs) address the latter by representing weights …
- They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
Alex Kwon · 7 juillet 2026
When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check. We treat the sender's communicat…
- NeSy-CSA: A Neuro-Symbolic Framework for Open-Ended Critical Scenario Attribution
Qitong Chu, Xunjie He, Chen Deng, Huaxin Pei, Yufeng Yue · 7 juillet 2026
Understanding why discovered scenarios become critical in scenario-based testing is essential for effectively leveraging them in decision-making systems. Reasoning about such criticality can be formulated as an attribution problem. However, across different decision-making tasks, the causes of criti…
- Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil · 7 juillet 2026
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identical verbal descriptions for identical inputs, and select de…
- Exploring the Rashomon Set for Concept-Based Models
Shihan Feng, Cheng Zhang, Michael Xi, Ethan Hsu, Lesia Semenova, Chudi Zhong · 7 juillet 2026
In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on fundamentally different internal logic. However, standard training procedures produce a single model, offering no practical way to explore alternatives that may be…
- InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
Yifan Luo, Zhennan Zhou, Bin Dong · 7 juillet 2026
Understanding the internal representations of large language models (LLMs) is a central challenge in interpretability research. Existing feature interpretability methods often rely on strong structural assumptions--such as linearity or sparsity--that may not hold in practice. In this work, we introd…
- Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators
Andrew Zhang, Chengzhan Li · 7 juillet 2026
Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscure the diagnostic question developers need most: which action changed the state in a useful direction? We introduce Agent Step Value (ASV), a state-transition me…
- Measuring What Matters: A Unified Evaluation Framework for GNN Explainability
Francesco Paolo Nerini, Mirko Zaffaroni, Paolo Baracco, Gabriele Ciravegna, Alan Perotti · 7 juillet 2026
Graph eXplainable AI (G-XAI) is increasingly important for making Graph Neural Networks interpretable and accountable. While a growing number of explainers are available, choosing the right method and assessing the trustworthiness of its outputs remains unclear. Consistent evaluation practices and a…
- When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
Peiying Zhu, Sidi Chang · 7 juillet 2026
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable. We study this problem in a hotel-pricing simulator where an agentic policy editor receives only region-level diagnostic…
- Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Johannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi Sugiyama · 7 juillet 2026
Reinforcement Learning from Human Feedback (RLHF) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Models (LMs). A common problem is reward hacking, where the policy may exploit inaccuracies of the reward and learn an unintended behavior. Most previous works add…
- Validation-Induced Shapley Shifts: How Validation Structure Distorts Data Valuation
Yinan Shen, Ziao Yang, Hongfu Liu · 7 juillet 2026
Shapley values are widely used to attribute value to training data based on their marginal contribution to performance on a validation set. Existing practice often assumes these values are stable once the training data and model are fixed. In this work, we uncover a systematic vulnerability: even mo…
- Out-of-Distribution Generalization of Risk Aversion in Language Models
Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie, Andrew Lin, Abhitej Bokka, Elliott Thornley · 7 juillet 2026
Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-reward strategies like cooperation over high-risk, high-reward strategies like rebellion, limiting the downsides of any misa…
- Unsupervised Features Mining via Activation Geometry
Amit LeVi, Elad David, Max Fomin · 7 juillet 2026
Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with labeled examples of a human-defined concept that may reflect human biases, and then identify how that concept is represented within the model, for example in its acti…
- Personalized Causal Recourse: A Human-In-The-Loop Approach
Denise Tampieri, Giovanni De Toni, Paolo Giudici · 7 juillet 2026
Algorithmic recourse addresses the challenge of providing tailored recommendations to users affected by unfavorable machine learning decisions, in potentially high-stakes scenarios. Traditional approaches to recourse often rely on the closest counterfactual explanations or assume a priori knowledge …
- Position: Use Sparse Autoencoders to Discover Unknowns
Kenny Peng, Rajiv Movva, Jon Kleinberg, Emma Pierson, Nikhil Garg · 7 juillet 2026
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptual distinction that reconciles competing narratives surrounding SAEs. We argue that even if SAEs may be less effective fo…
- Explainable Novel Category Discovery in Semantic Concept Space
Ifrat Ikhtear Uddin, Yang Zhou, KC Santosh, Longwei Wang · 7 juillet 2026
Novel category discovery aims to identify unseen classes from unlabeled data by transferring knowledge from labeled categories, but most existing methods perform discovery in opaque latent feature spaces. As a result, they may separate novel categories accurately while providing little insight into …
- Decentralized Aggregation of LLM Predictions via Wagering Mechanisms
Yuhong Luo, David M. Pennock, Xintong Wang · 7 juillet 2026
It is increasingly common to aggregate predictions from multiple LLMs, each with domain expertise or access to private tools and data, to improve collective prediction performance. In decentralized settings, aggregation weights need to be determined without access to models' private information and …
- Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
Akchunya Chanchal, David A. Kelly, Hana Chockler · 7 juillet 2026
Perturbation-based explainability methods face criticism due to their reliance on out-of-distribution mutants. This raises doubts about the quality of the explanations. In this paper, we introduce a novel forward pass paradigm, Activation-Deactivation (AD), which obviates the need for perturbation o…
- Faithfulness to Refusal: A Causal Audit of Neuron Selectors
Ananth Eswar, Pratinav Seth, Utsav Avaiya, Vinay Kumar Sankarapu · 7 juillet 2026
Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly. We address this with two paired audits built on one-shot neur…
- Conditional Inference Trees and Forests for Feature Selection
Robert Milletich, Justin Downes, Steve Goley, Newel Hirst · 3 juillet 2026
Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive. We study CIT and CIF as top-$k$ feature-…
- Adaptive Group-Based Counterfactual Explanations for Time-Series Rehabilitation Data
Emmanuel C. Chukwu, Rianne M. Schouten, Monique Tabak, Mykola Pechenizkiy · 3 juillet 2026
Counterfactual explanations (CEs) for multivariate time-series classifiers are often difficult to interpret in domains where experts reason in terms of semantic feature groups rather than individual channels. In rehabilitation movement analysis with multi-sensor inertial measurement units (IMUs), cl…
- Bi-NAS: Towards Effective and Personalized Explanation for Recommender Systems via Bi-Level Neural Architecture Search
Longfeng Wu, Yao Zhou, Tong Zeng, Zhimin Peng, Bhanu Pratap Singh Rawat, Lecheng Zheng, Giovanni Seni, Dawei Zhou · 3 juillet 2026
Recommender systems are vital in helping users navigate vast amounts of information, offering personalized suggestions and effective explanations for these recommendations. While previous efforts have attempted to provide such explanations, evaluating their effectiveness across various scenarios rem…
