Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 222 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Gradient Boosted Risk Scores
Costa Georgantas, Jonas Richiardi · 5 mai 2026
Risk scores are an interpretable and actionable class of machine learning models with applications in medicine, insurance, and risk management. Unlike most computational methods, risk scores are designed to be computed by a human by attributing points to a data sample based on a limited set of crite…
- Binary Rewards and Reinforcement Learning: Fundamental Challenges
Marc Dymetman · 5 mai 2026
Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for improving reasoning in language models, yet models trained with RLVR often suffer from diversity collapse: while single-sample accuracy improves, multi-sample coverage degrades, sometimes falling below the base …
- Model Routing as a Trust Problem: Route Receipts for Adaptive AI Systems
Vincent Schmalbach · 5 mai 2026
AI products often route requests through version aliases, service tiers, tool choices, regional endpoints, fallback rules, or safety handling before responding. These routing steps are documented product surfaces in several widely used AI platforms and serving stacks. Routing helps AI services sta…
- ParaRNN: An Interpretable and Parallelizable Recurrent Neural Network for Time-Dependent Data
Yuxi Cai, Lan Li, Feiqing Huang, Guodong Li · 5 mai 2026
The proliferation of large-scale and structurally complex data has spurred the integration of machine learning methods into statistical modeling. Recurrent neural networks (RNNs), a foundational class of models for time-dependent data, can be viewed as nonlinear extensions of classical autoregressiv…
- Metric Unreliability in Multimodal Machine Unlearning: A Systematic Analysis and Principled Unified Score
Abdullah Ahmad Khan, Hamid Laga, Ferdous Sohel · 5 mai 2026
Machine unlearning in Vision-Language Models (VLMs) is required for compliance with the General Data Protection Regulation (GDPR), yet current evaluation practices are inconsistent. We present the first systematic study of metric reliability in multimodal unlearning. Five standard metrics, Forget Ac…
- Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models
Inoussa Mouiche · 5 mai 2026
Preference optimization has become a central paradigm for aligning large language models with human feedback. Direct Preference Optimization (DPO) simplifies reinforcement learning from human feedback by directly optimizing pairwise preferences, removing the need for reward modeling and policy optim…
- On the explainability of max-plus neural networks
Ikhlas Enaieh (S2A, LTCI), Olivier Fercoq (S2A, LTCI), Garc\'ia \'Angel (DATSI, UPM) · 5 mai 2026
We investigate the explanability properties of the recently proposed linear-min-max neural networks. At initialization, they can be interpreted as k-medoids with the infinity norm as a distance. Then, they are trained using subgradient descent to better fit the data. The model has been shown to be a…
- Constructing Interpretable Features from Compositional Neuron Groups
Or Shafran, Atticus Geiger, Mor Geva · 5 mai 2026
A central goal for mechanistic interpretability has been to identify the right units of analysis in large language models (LLMs) that causally explain their outputs. While early work focused on individual neurons, evidence that neurons often encode multiple concepts has motivated a shift toward anal…
- Synthetic Designed Experiments for Diagnosing Vision Model Failure
Krisanu Sarkar · 5 mai 2026
Current synthetic data pipelines for computer vision generate images without diagnosing what the downstream model actually needs. This open-loop paradigm treats synthetic data as cheap real data, randomly sampling the generator's output space and hoping to cover the model's failure modes. We argue t…
- $\phi$-Table: A Statistical Explanation for Global SHAP
Dongseok Kim, Hyoungsun Choi, Mohamed Jismy Aashik Rasool, Gisung Oh · 5 mai 2026
Global SHAP explanations are typically presented as feature-importance rankings, which identify variables that matter to a black-box model but do not indicate whether their effects admit clear directional summaries, how uncertain those summaries are, or how faithfully they represent the fitted respo…
- How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
Daniel da Silva Costa, Pedro Nuno de Souza Moura, Adriana C. F. Alvim · 5 mai 2026
In recent years, several advances have been observed in Deep Learning with surprising results. Models in this area have been increasingly used in numerous applications, including those sensitive to human life, which require clear explanations and justifications. Various explainability methods have b…
- Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling
Brice Valentin Kok-Shun, Johnny Chan, Gabrielle Peko, David Sundaram · 5 mai 2026
Agentopic is a novel agent-based workflow for explainable topic modeling that leverages the reasoning capabilities of Large Language Models (LLMs). Existing topic modeling approaches such as Latent Dirichlet Allocation (LDA) and BERTopic often lack transparency on how topics are assigned or grouped.…
- Retrieval-Augmented Reasoning for Chartered Accountancy
Jatin Gupta, Akhil Sharma, Saransh Singhania, Ali Imam Abidi · 4 mai 2026
The inception of Large Language Models (LLMs) has catalyzed AI adoption in the finance sector, yet their reliability in complex, jurisdiction-specific tasks like Indian Chartered Accountancy (CA) remains limited. The models display difficulty in executing numerical tasks which require multiple steps…
- OTSS: Output-Targeted Soft Segmentation for Contextual Decision-Weight Learning
Renjun Hu, Hyun-Soo Ahn · 4 mai 2026
Many machine learning systems make constrained decisions by optimizing factorized objectives, but the context-specific objective is often treated as fixed. We study contextual decision-weight learning: from logged decisions and proxy outputs, learn an optimizer-facing weight vector w(x) over interpr…
- Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions
Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu · 4 mai 2026
Large language models (LLMs) are increasingly tasked with strategic decision-making under incomplete information, such as in negotiation and policymaking. While LLMs can excel at many such tasks, they also fail in ways that are poorly understood. We shed light on these failures by uncovering two fun…
- Position: agentic AI orchestration should be Bayes-consistent
Theodore Papamarkou, Pierre Alquier, Matthias Bauer, Wray Buntine, Andrew Davison, Gintare Karolina Dziugaite, Maurizio Filippone, Andrew Y. K. Foong, Vincent Fortuin, Dimitris Fouskakis, Jes Frellsen, Eyke H\"ullermeier, Theofanis Karaletsos, Mohammad Emtiyaz Khan, Nikita Kotelevskii, Salem Lahlou, Yingzhen Li, Fang Liu, Clare Lyle, Thomas M\"ollenhoff, Konstantina Palla, Maxim Panov, Yusuf Sale, Kajetan Schweighofer, Artem Shelmanov, Siddharth Swaroop, Martin Trapp, Willem Waegeman, Andrew Gordon Wilson, Alexey Zaytsev · 4 mai 2026
LLMs excel at predictive tasks and complex reasoning tasks, but many high-value deployments rely on decisions under uncertainty, for example, which tool to call, which expert to consult, or how many resources to invest. While the usefulness and feasibility of Bayesian approaches remain unclear for L…
- Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework
Sheza Munir, Ahanaf Rodoshi, Sumin Lee, Feiran Chang, Xujie Si, Syed Ishtiaque Ahmed · 4 mai 2026
Standard methods for aggregating natural language judgments, such as majority voting, often fail to produce logically consistent results when applied to high-conflict domains, treating differing opinions as noise. We propose a neuro-symbolic aggregation framework that formalizes conflict resolution …
- Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
Yuxuan Gao, Megan Wang, Yi Ling Yu · 4 mai 2026
Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) tuple at which a specific quantization, decoding strategy, region, and serving stack is exposed. We …
- Binary Spiking Neural Networks as Causal Models
Aditya Kar (CNRS, IRIT), Emiliano Lorini (CNRS, IRIT), Timoth\'ee Masquelier (CNRS, CERCO UMR5549) · 1 mai 2026
We provide a causal analysis of Binary Spiking Neural Networks (BSNNs) to explain their behavior. We formally define a BSNN and represent its spiking activity as a binary causal model. Thanks to this causal representation, we are able to explain the output of the network by leveraging logic-based me…
- Epistemic reflections on AI answering our questions: overwatch, erudite, logician, interlocutor
Johan F. Hoorn, Ella-Jenna Oosterglorenwoud · 1 mai 2026
Currently, there is a trend for the wider public to rely on LLMs for financial or legal consultation, medical and mental support (Chatterji et al., 2025), often accepting the advice provided without necessarily seeking logical verification or empirical validation. While one might be fortunate enough…
- Learning-to-Explain through 20Q Gaming: An Explainable Recommender for Cybersecurity Education
Mary Nusrat, Sarfuddin Bhuiyan, Gahangir Hossain · 1 mai 2026
The growing sophistication of contemporary cyber threats necessitates a more effective and adaptive approach to cybersecurity training. Intuitive and adaptive approaches to learning, which are often required, are not provided in traditional learning methods. In this article, we present a new educati…
- Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective
Ziyao Xu, Cong Wang, Houfeng Wang · 1 mai 2026
Compositional generalization tests are often used to estimate the compositionality of LLMs. However, such tests have the following limitations: (1) they only focus on the output results without considering LLMs' understanding of sample compositionality, resulting in explainability defects; (2) they …
- Do Sparse Autoencoders Capture Concept Manifolds?
Usha Bhalla, Thomas Fel, Can Rager, Sheridan Feucht, Tal Haklay, Daniel Wurgaft, Siddharth Boppana, Matthew Kowal, Vasudev Shyam, Jack Merullo, Atticus Geiger, Ekdeep Singh Lubana · 1 mai 2026
Sparse autoencoders (SAEs) are widely used to extract interpretable features from neural network representations, often under the implicit assumption that concepts correspond to independent linear directions. However, a growing body of evidence suggests that many concepts are instead organized along…
- DDO-RM: Distribution-Level Policy Improvement after Reward Learning
Tiantian Zhang, Jierui Zuo, Michael Chen, Wenping Wang · 1 mai 2026
Recent theory suggests that reward-model-first methods can be more sample-efficient than direct policy fitting when the reward function is statistically simpler than the induced policy. We propose DDO-RM, a finite-candidate decision-optimization method that converts reward scores into an explicit ta…
- CoAX: Cognitive-Oriented Attribution eXplanation User Model of Human Understanding of AI Explanations
Louth Bin Rawshan, Zhuoyu Wang, Brian Y. Lim · 1 mai 2026
Explainable AI (XAI) aims to improve user understanding and decisions when using AI models. However, despite innovations in XAI, recent user evaluations reveal that this goal remains elusive. Understanding human cognition can help explain why users struggle to effectively use AI explanations. Focusi…
