Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 226 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Dynamic Feature Selection based on Rule-based Learning for Explainable Classification with Uncertainty Quantification
Javier Fumanal-Idocin, Raquel Fernandez-Peralta, Javier Andreu-Perez · 3 décembre 2025
Dynamic feature selection (DFS) offers a compelling alternative to traditional, static feature selection by adapting the selected features to each individual sample. This provides insights into the decision-making process for each case, which makes DFS especially significant in settings where decisi…
- Computational Copyright: Towards A Royalty Model for Music Generative AI
Junwei Deng, Xirui Jiang, Shiyuan Zhang, Shichang Zhang, Himabindu Lakkaraju, Ruijiang Gao, Chris Donahue, Jiaqi W. Ma · 3 décembre 2025
The rapid rise of generative AI has intensified copyright and economic tensions in creative industries, particularly in music. Current approaches addressing this challenge often focus on preventing infringement or establishing one-time licensing, which fail to provide the sustainable, recurring econ…
- A Framework for Causal Concept-based Model Explanations
Anna Rodum Bj{\o}ru, Jacob Lysn{\ae}s-Larsen, Oskar J{\o}rgensen, Inga Str\"umke, Helge Langseth · 3 décembre 2025
This work presents a conceptual framework for causal concept-based post-hoc Explainable Artificial Intelligence (XAI), based on the requirements that explanations for non-interpretable models should be understandable as well as faithful to the model being explained. Local and global explanations are…
- The Effect of Enforcing Fairness on Reshaping Explanations in Machine Learning Models
Joshua Wolff Anderson, Shyam Visweswaran · 3 décembre 2025
Trustworthy machine learning in healthcare requires strong predictive performance, fairness, and explanations. While it is known that improving fairness can affect predictive performance, little is known about how fairness improvements influence explainability, an essential ingredient for clinical t…
- Enforcing Orderedness to Improve Feature Consistency
Sophie L. Wang, Alex Quach, Nithin Parsan, John J. Yang · 3 décembre 2025
Sparse autoencoders (SAEs) have been widely used for interpretability of neural networks, but their learned features often vary across seeds and hyperparameter settings. We introduce Ordered Sparse Autoencoders (OSAE), which extend Matryoshka SAEs by (1) establishing a strict ordering of latent feat…
- STRIDE: A Systematic Framework for Selecting AI Modalities - Agentic AI, AI Assistants, or LLM Calls
Shubhi Asthana, Bing Zhang, Chad DeLuca, Ruchi Mahindru, Hima Patel · 3 décembre 2025
The rapid shift from stateless large language models (LLMs) to autonomous, goal-driven agents raises a central question: When is agentic AI truly necessary? While agents enable multi-step reasoning, persistent memory, and tool orchestration, deploying them indiscriminately leads to higher cost, comp…
- ZIP-RC: Zero-overhead Inference-time Prediction of Reward and Cost for Adaptive and Interpretable Generation
Rohin Manvi, Joey Hong, Tim Seyde, Maxime Labonne, Mathias Lechner, Sergey Levine · 2 décembre 2025
Large language models excel at reasoning but lack key aspects of introspection, including anticipating their own success and the computation required to achieve it. Humans use real-time introspection to decide how much effort to invest, when to make multiple attempts, when to stop, and when to signa…
- Faster Verified Explanations for Neural Networks
Alessandro De Palma, Greta Dolcetti, Caterina Urban · 2 décembre 2025
Verified explanations are a theoretically-principled way to explain the decisions taken by neural networks, which are otherwise black-box in nature. However, these techniques face significant scalability challenges, as they require multiple calls to neural network verifiers, each of them with an exp…
- AlignSAE: Concept-Aligned Sparse Autoencoders
Minglai Yang, Xinyu Guo, Mihai Surdeanu, Liangming Pan · 2 décembre 2025
Large Language Models (LLMs) encode factual knowledge within hidden parametric spaces that are difficult to inspect or control. While Sparse Autoencoders (SAEs) can decompose hidden activations into more fine-grained, interpretable features, they often struggle to reliably align these features with …
- Multiclass threshold-based classification
Francesco Marchetti, Edoardo Legnaro, Sabrina Guastavino · 2 décembre 2025
In this paper, we introduce a threshold-based framework for multiclass classification that generalizes the standard argmax rule. This is done by replacing the probabilistic interpretation of softmax outputs with a geometric one on the multidimensional simplex, where the classification depends on a m…
- One Swallow Does Not Make a Summer: Understanding Semantic Structures in Embedding Spaces
Yandong Sun, Qiang Huang, Ziwei Xu, Yiqun Sun, Yixuan Tang, Anthony K. H. Tung · 2 décembre 2025
Embedding spaces are fundamental to modern AI, translating raw data into high-dimensional vectors that encode rich semantic relationships. Yet, their internal structures remain opaque, with existing approaches often sacrificing semantic coherence for structural regularity or incurring high computati…
- EXP-CAM: Explanation Generation and Circuit Discovery Using Classifier Activation Matching
Pirzada Suhail, Aditya Anand, Amit Sethi · 2 décembre 2025
Machine learning models, by virtue of training, learn a large repertoire of decision rules for any given input, and any one of these may suffice to justify a prediction. However, in high-dimensional input spaces, such rules are difficult to identify and interpret. In this paper, we introduce EXP-CAM…
- Unsupervised decoding of encoded reasoning using language model interpretability
Ching Fang, Samuel Marks · 2 décembre 2025
As large language models become increasingly capable, there is growing concern that they may develop reasoning processes that are encoded or hidden from human oversight. To investigate whether current interpretability techniques can penetrate such encoded reasoning, we construct a controlled testbed…
- A Self-explainable Model of Long Time Series by Extracting Informative Structured Causal Patterns
Ziqian Wang, Yuxiao Cheng, Jinli Suo · 2 décembre 2025
Explainability is essential for neural networks that model long time series, yet most existing explainable AI methods only produce point-wise importance scores and fail to capture temporal structures such as trends, cycles, and regime changes. This limitation weakens human interpretability and trust…
- Probabilistic Neuro-Symbolic Reasoning for Sparse Historical Data: A Framework Integrating Bayesian Inference, Causal Models, and Game-Theoretic Allocation
Saba Kublashvili · 2 décembre 2025
Modeling historical events poses fundamental challenges for machine learning: extreme data scarcity (N << 100), heterogeneous and noisy measurements, missing counterfactuals, and the requirement for human interpretable explanations. We present HistoricalML, a probabilistic neuro-symbolic framework t…
- The Impact of Concept Explanations and Interventions on Human-Machine Collaboration
Jack Furby, Dan Cunnington, Dave Braines, Alun Preece · 2 décembre 2025
Deep Neural Networks (DNNs) are often considered black boxes due to their opaque decision-making processes. To reduce their opacity Concept Models (CMs), such as Concept Bottleneck Models (CBMs), were introduced to predict human-defined concepts as an intermediate step before predicting task labels.…
- Decision Tree Embedding by Leaf-Means
Cencheng Shen, Yuexiao Dong, Carey E. Priebe · 2 décembre 2025
Decision trees and random forest remain highly competitive for classification on medium-sized, standard datasets due to their robustness, minimal preprocessing requirements, and interpretability. However, a single tree suffers from high estimation variance, while large ensembles reduce this variance…
- Energy-Aware Data-Driven Model Selection in LLM-Orchestrated AI Systems
Daria Smirnova, Hamid Nasiri, Marta Adamska, Zhengxin Yu, Peter Garraghan · 2 décembre 2025
As modern artificial intelligence (AI) systems become more advanced and capable, they can leverage a wide range of tools and models to perform complex tasks. Today, the task of orchestrating these models is often performed by Large Language Models (LLMs) that rely on qualitative descriptions of mode…
- Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
Saeed AlMarri, Mathieu Ravaut, Kristof Juhasz, Gautier Marti, Hamdan Al Ahbabi, Ibrahim Elfadel · 2 décembre 2025
Large Language Models (LLMs) have attracted significant attention for classification tasks, offering a flexible alternative to trusted classical machine learning models like LightGBM through zero-shot prompting. However, their reliability for structured tabular data remains unclear, particularly in …
- Probing the "Psyche'' of Large Reasoning Models: Understanding Through a Human Lens
Yuxiang Chen, Zuohan Wu, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen · 2 décembre 2025
Large reasoning models (LRMs) have garnered significant attention from researchers owing to their exceptional capability in addressing complex tasks. Motivated by the observed human-like behaviors in their reasoning processes, this paper introduces a comprehensive taxonomy to characterize atomic rea…
- Pushing the Boundaries of Interpretability: Incremental Enhancements to the Explainable Boosting Machine
Isara Liyanage, Uthayasanker Thayasivam · 2 décembre 2025
The widespread adoption of complex machine learning models in high-stakes domains has brought the "black-box" problem to the forefront of responsible AI research. This paper aims at addressing this issue by improving the Explainable Boosting Machine (EBM), a state-of-the-art glassbox model that deli…
- CRAwDAD: Causal Reasoning Augmentation with Dual-Agent Debate
Finn G. Vamosi, Nils D. Forkert · 1 décembre 2025
When people reason about cause and effect, they often consider many competing "what if" scenarios before deciding which explanation fits best. Analogously, advanced language models capable of causal inference can consider multiple interventions and counterfactuals to judge the validity of causal cla…
- On the Effect of Regularization on Nonparametric Mean-Variance Regression
Eliot Wong-Toi, Alex Boyd, Vincent Fortuin, Stephan Mandt · 1 décembre 2025
Uncertainty quantification is vital for decision-making and risk assessment in machine learning. Mean-variance regression models, which predict both a mean and residual noise for each data point, provide a simple approach to uncertainty quantification. However, overparameterized mean-variance models…
- TabPFN: One Model to Rule Them All?
Qiong Zhang, Yan Shuo Tan, Qinglong Tian, Pengfei Li · 1 décembre 2025
Hollmann et al. (Nature 637 (2025) 319-326) recently introduced TabPFN, a transformer-based deep learning model for regression and classification on tabular data, which they claim "outperforms all previous methods on datasets with up to 10,000 samples by a wide margin, using substantially less train…
- LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations
Payal Varshney, Adriano Lucieri, Christoph Balada, Sheraz Ahmed, Andreas Dengel · 1 décembre 2025
Video-based AI systems are increasingly adopted in safety-critical domains such as autonomous driving and healthcare. However, interpreting their decisions remains challenging due to the inherent spatiotemporal complexity of video data and the opacity of deep learning models. Existing explanation te…
