Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 222 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- A Novel Approach to Explainable AI with Quantized Active Ingredients in Decision Making
A. M. A. S. D. Alagiyawanna, Asoka Karunananda, Thushari Silva, A. Mahasinghe · 14 janvier 2026
Artificial Intelligence (AI) systems have shown good success at classifying. However, the lack of explainability is a true and significant challenge, especially in high-stakes domains, such as health and finance, where understanding is paramount. We propose a new solution to this challenge: an expla…
- Bridging the Trust Gap: Clinician-Validated Hybrid Explainable AI for Maternal Health Risk Assessment in Bangladesh
Farjana Yesmin, Nusrat Shirmin, Suraiya Shabnam Bristy · 14 janvier 2026
While machine learning shows promise for maternal health risk prediction, clinical adoption in resource-constrained settings faces a critical barrier: lack of explainability and trust. This study presents a hybrid explainable AI (XAI) framework combining ante-hoc fuzzy logic with post-hoc SHAP expla…
- YaPO: Learnable Sparse Activation Steering Vectors for Domain Adaptation
Abdelaziz Bounhar, Rania Hossam Elmohamady Elbadry, Hadi Abdine, Preslav Nakov, Michalis Vazirgiannis, Guokan Shang · 14 janvier 2026
Steering Large Language Models (LLMs) through activation interventions has emerged as a lightweight alternative to fine-tuning for alignment and personalization. Recent work on Bi-directional Preference Optimization (BiPO) shows that dense steering vectors can be learned directly from preference dat…
- Self-Certification of High-Risk AI Systems: The Example of AI-based Facial Emotion Recognition
Gregor Autischer, Kerstin Waxnegger, Dominik Kowald · 14 janvier 2026
The European Union's Artificial Intelligence Act establishes comprehensive requirements for high-risk AI systems, yet the harmonized standards necessary for demonstrating compliance remain not fully developed. In this paper, we investigate the practical application of the Fraunhofer AI assessment ca…
- TabPFN Through The Looking Glass: An interpretability study of TabPFN and its internal representations
Aviral Gupta, Armaan Sethi, Dhruv Kumar · 14 janvier 2026
Tabular foundational models are pre-trained models designed for a wide range of tabular data tasks. They have shown strong performance across domains, yet their internal representations and learned concepts remain poorly understood. This lack of interpretability makes it important to study how these…
- Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
Kaivalya Rawal, Eoin Delaney, Zihao Fu, Sandra Wachter, Chris Russell · 14 janvier 2026
Explainable artificial intelligence (XAI) is concerned with producing explanations indicating the inner workings of models. For a Rashomon set of similarly performing models, explanations provide a way of disambiguating the behavior of individual models, helping select models for deployment. However…
- EviNAM: Intelligibility and Uncertainty via Evidential Neural Additive Models
S\"oren Schleibaum, Anton Frederik Thielmann, Julian Teusch, Benjamin S\"afken, J\"org P. M\"uller · 14 janvier 2026
Intelligibility and accurate uncertainty estimation are crucial for reliable decision-making. In this paper, we propose EviNAM, an extension of evidential learning that integrates the interpretability of Neural Additive Models (NAMs) with principled uncertainty estimation. Unlike standard Bayesian n…
- Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations
Myra Cheng, Vinodkumar Prabhakaran, Alice Oh, Hayk Stepanyan, Aishwarya Verma, Charu Kalia, Erin MacMurray van Liemt, Sunipa Dev · 14 janvier 2026
Generative AI models ought to be useful and safe across cross-cultural contexts. One critical step toward this goal is understanding how AI models adhere to sociocultural norms. While this challenge has gained attention in NLP, existing work lacks both nuance and coverage in understanding and evalua…
- Forward versus Backward: Comparing Reasoning Objectives in Direct Preference Optimization
Murtaza Nikzad, Raghuram Ramanujan · 13 janvier 2026
Large language models exhibit impressive reasoning capabilities yet frequently generate plausible but incorrect solutions, a phenomenon commonly termed hallucination. This paper investigates the effect of training objective composition on reasoning reliability through Direct Preference Optimization.…
- Explaining Machine Learning Predictive Models through Conditional Expectation Methods
Silvia Ruiz-Espa\~na (ITI, Universitat Polit\`ecnica de Val\`encia, Val\`encia, Spain), Laura Arnal (ITI, Universitat Polit\`ecnica de Val\`encia, Val\`encia, Spain), Fran\c{c}ois Signol (ITI, Universitat Polit\`ecnica de Val\`encia, Val\`encia, Spain), Juan-Carlos Perez-Cortes (ITI, Universitat Polit\`ecnica de Val\`encia, Val\`encia, Spain), Joaquim Arlandis (ITI, Universitat Polit\`ecnica de Val\`encia, Val\`encia, Spain) · 13 janvier 2026
The rapid adoption of complex Artificial Intelligence (AI) and Machine Learning (ML) models has led to their characterization as black boxes due to the difficulty of explaining their internal decision-making processes. This lack of transparency hinders users' ability to understand, validate and trus…
- PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
Ruiyi Ding, Yongxuan Lv, Xianhui Meng, Jiahe Song, Chao Wang, Chen Jiang, Yuan Cheng · 13 janvier 2026
Policy optimization for large language models often suffers from sparse reward signals in multi-step reasoning tasks. Critic-free methods like GRPO assign a single normalized outcome reward to all tokens, providing limited guidance for intermediate reasoning . While Process Reward Models (PRMs) offe…
- Towards Public Administration Research Based on Interpretable Machine Learning
Zhanyu Liu, Yang Yu · 13 janvier 2026
Causal relationships play a pivotal role in research within the field of public administration. Ensuring reliable causal inference requires validating the predictability of these relationships, which is a crucial precondition. However, prediction has not garnered adequate attention within the realm …
- Hidden Monotonicity: Explaining Deep Neural Networks via their DC Decomposition
Jakob Paul Zimmermann, Georg Loho · 13 janvier 2026
It has been demonstrated in various contexts that monotonicity leads to better explainability in neural networks. However, not every function can be well approximated by a monotone neural network. We demonstrate that monotonicity can still be used in two ways to boost explainability. First, we use a…
- Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity
Hongjun An, Yiliang Song, Jiangan Chen, Jiawei Shao, Chi Zhang, Xuelong Li · 13 janvier 2026
Large Language Model (LLM) training often optimizes for preference alignment, rewarding outputs that are perceived as helpful and interaction-friendly. However, this preference-oriented objective can be exploited: manipulative prompts can steer responses toward user-appeasing agreement and away from…
- Bag of Coins: A Statistical Probe into Neural Confidence Structures
Agnideep Aich, Sameera Hewage, Md Monzur Murshed, Bruce Wade, Ashit Baran Aich · 13 janvier 2026
Modern neural networks often produce miscalibrated confidence scores and struggle to detect out-of-distribution (OOD) inputs, while most existing methods post-process outputs without testing internal consistency. We introduce the Bag-of-Coins (BoC) probe, a non-parametric diagnostic of logit coheren…
- Beyond Perfect Scores: Proof-by-Contradiction for Trustworthy Machine Learning
Dushan N. Wadduwage, Dineth Jayakody, Leonidas Zimianitis · 13 janvier 2026
Machine learning (ML) models show strong promise for new biomedical prediction tasks, but concerns about trustworthiness have hindered their clinical adoption. In particular, it is often unclear whether a model relies on true clinical cues or on spurious hierarchical correlations in the data. This p…
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
Joshua Ashkinaze, Hua Shen, Saipranav Avula, Eric Gilbert, Ceren Budak · 13 janvier 2026
We introduce the Deep Value Benchmark (DVB), an evaluation framework that directly tests whether large language models (LLMs) learn fundamental human values or merely surface-level preferences. This distinction is critical for AI alignment: Systems that capture deeper values are likely to generalize…
- From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards
Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos · 13 janvier 2026
Explainable AI (XAI) in high-stakes domains should help stakeholders trust and verify system outputs. Yet Chain-of-Thought methods reason before concluding, and logical gaps or hallucinations can yield conclusions that do not reliably align with their rationale. Thus, we propose "Result -> Justify",…
- ConSensus: Multi-Agent Collaboration for Multimodal Sensing
Hyungjun Yoon, Mohammad Malekzadeh, Sung-Ju Lee, Fahim Kawsar, Lorena Qendro · 13 janvier 2026
Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However, accurately interpreting heterogeneous multimodal sensor data remains a fundamental challenge. We show that a single monolithic LLM often fails to reaso…
- Why are there many equally good models? An Anatomy of the Rashomon Effect
Harsh Parikh · 13 janvier 2026
The Rashomon effect -- the existence of multiple, distinct models that achieve nearly equivalent predictive performance -- has emerged as a fundamental phenomenon in modern machine learning and statistics. In this paper, we explore the causes underlying the Rashomon effect, organizing them into thre…
- Explainability of Complex AI Models with Correlation Impact Ratio
Poushali Sengupta, Rabindra Khadka, Sabita Maharjan, Frank Eliassen, Yan Zhang, Shashi Raj Pandey, Pedro G. Lind, Anis Yazidi · 13 janvier 2026
Complex AI systems make better predictions but often lack transparency, limiting trustworthiness, interpretability, and safe deployment. Common post hoc AI explainers, such as LIME, SHAP, HSIC, and SAGE, are model agnostic but are too restricted in one significant regard: they tend to misrank correl…
- Reasoning Models Will Blatantly Lie About Their Reasoning
William Walden · 13 janvier 2026
It has been shown that Large Reasoning Models (LRMs) may not *say what they think*: they do not always volunteer information about how certain parts of the input influence their reasoning. But it is one thing for a model to *omit* such information and another, worse thing to *lie* about it. Here, we…
- Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation
Hua Ye, Siyuan Chen, Ziqi Zhong, Canran Xiao, Haoliang Zhang, Yuhan Wu, Fei Shen · 13 janvier 2026
Large language models (LLMs) equipped with retrieval--the Retrieval-Augmented Generation (RAG) paradigm--should combine their parametric knowledge with external evidence, yet in practice they often hallucinate, over-trust noisy snippets, or ignore vital context. We introduce TCR (Transparent Conflic…
- SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis
Zihao Fu, Xufeng Duan, Zhenguang G. Cai · 13 janvier 2026
Large language models excel across diverse domains, yet their deployment in healthcare, legal systems, and autonomous decision-making remains limited by incomplete understanding of their internal mechanisms. As these models integrate into high-stakes systems, understanding how they encode capabiliti…
- Covariance-Driven Regression Trees: Reducing Overfitting in CART
Likun Zhang, Wei Ma · 13 janvier 2026
Decision trees are powerful machine learning algorithms, widely used in fields such as economics and medicine for their simplicity and interpretability. However, decision trees such as CART are prone to overfitting, especially when grown deep or the sample size is small. Conventional methods to redu…
