Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 219 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges
Swastik Roy, Rajkumar Pujari, Tharindu Kumarage, Charith Peris, Rahul Gupta, Anna Rumshisky, Pradeep Natarajan, Venkatesh Saligrama · 1 juin 2026
LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be ``helpful and factual'' can reward polished answers that invent facts or violate user intent. We treat reusable rubrics a…
- Learning to Perceive the World Through Control: Empowerment-Based Representation Learning
Mahsa Bastankhah, Sophie Broderick, Benjamin Eysenbach · 1 juin 2026
In many practical reinforcement learning environments, observations are far higher-dimensional than the variables that matter for control. In this work, we ask: can we learn representations that capture only control-relevant features of the environment? We study this question through the empowerment…
- Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
Kyle Moore, Jesse Roberts, Daryl Watson · 1 juin 2026
There has been much recent interest in evaluating large language models for uncertainty calibration to facilitate model control and modulate user trust. Inference time uncertainty, which may provide a real-time signal to the model or external control modules, is particularly important for applying t…
- LLMs Without Deep Neural Networks: New Architecture, Benefits and Case Study
Vincent Granville · 1 juin 2026
The purpose of this article is to provide validation to my deep neural network alternative in the context of LLMs. Very recently, there has been a significant interest by Chinese researchers in a model called RBF network, as a substitute to standard DNNs, with increased explainability and higher acc…
- Beyond Additive Decompositions: Interpretability Through Separability
Jinyang Liu, Munir Eberhardt Hiabu · 1 juin 2026
Interpretable machine learning requires models that are accurate and structurally faithful to the data.Existing explainability methods rely heavily on additive representations (e.g., Generalized Additive Models (GAMs), SHapley Additive exPlanations (SHAP), functional ANOVA), which can suffer from si…
- Algorithmic Recourse of In-Context Learning for Tabular Data
Wenshuo Dong, Jiaming Zhang, Shaopneg Fu, Hongbin Lin, Di Wang, Lijie Hu · 1 juin 2026
As predictive models are increasingly deployed in high-stakes settings such as credit approval, there is a growing need for post-hoc methods that provide recourse to affected individuals. Many such models operate on tabular data, where features correspond to real-world attributes. Recently, in-conte…
- TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories
Junjie Nian, Kang Chen, Ge Zhang, Yixin Cao, Yugang Jiang · 1 juin 2026
Agent benchmarks increasingly record rich interaction trajectories, yet evaluation often reduces each rollout to a pass rate or reward score. We introduce TraceGraph, a graph-based framework that turns released multi-model agent trajectories into shared decision landscapes. For each task, TraceGraph…
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Iv\'an Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, Arthur Conmy · 1 juin 2026
Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized reasoning can give an incorrect picture of how models arrive at conclusions (unfaithfulness). In this work, we show tha…
- FlagGAM: Rule-Based Generalized Additive Modeling for Explainable Tabular Prediction
Zijie Zhao, Roy E. Welsch · 1 juin 2026
Tabular prediction in high-stakes domains requires models that are accurate, transparent, and robust to imperfect inputs. We propose FlagGAM, a rule-defined basis framework that separates feature-level rule construction from prediction. A Flag Core Module converts numerical and categorical variables…
- When Are Multimodal Predictions Biologically Supported? A Diagnostic Evaluation Framework
Dylan Steiner, Gustavo Arango-Argoty, Gerald Sun, Etai Jacob · 1 juin 2026
Multimodal models in oncology can produce accurate predictions, but accurate prediction does not reveal whether the model has learned biology that is shared across modalities, biology confined to one modality, or spurious correlations that reflect confounders rather than genuine biology. We introduc…
- Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance
Do\u{g}ukan Ba\u{g}c{\i}, Bernt Schiele, Simone Schaub-Meyer, Jonas Fischer, Robin Hesse · 1 juin 2026
Deep neural networks (DNNs) are widely used, but interpreting what they actually learn remains difficult. A major obstacle is that individual neurons often encode multiple unrelated concepts, obscuring the decision process of the network. While prior work, such as sparse autoencoders, can separate t…
- Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution
Tarun Kota · 1 juin 2026
Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing oracle systems tradeoff fast but brittle automation against accurate but costly human arbitration. Single-LLM oracles achieve meaningful accuracy but …
- Density-Guided Robust Counterfactual Explanations on Tabular Data under Model Multiplicity
Jun Tan, Qing Guo, Zicheng Xu, Jinglin Li, Qi Fang, Ning Gui · 1 juin 2026
Counterfactual explanations (CEs) are essential for actionable recourse, yet their reliability is often compromised in low-density regions, where classifiers exhibit high variance. Unlike existing methods that rely on expensive ensemble intersections to define stability, we propose \textit{DensityFl…
- TUX: Measuring Human--AI Tacit Understanding
Yueshen Li, Hanyi Min, Vedant Das Swain, Koustuv Saha · 1 juin 2026
As large language models (LLMs) increasingly act as collaborative partners, human--AI alignment is often evaluated through explicit task success, accuracy, or reward optimization. Yet many collaborative settings depend on tacit understanding: whether an agent can align with a human's evaluative stan…
- Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation
Gr\'egoire Martinon, Ibrahim Merad, Mohammed Raki · 1 juin 2026
Reliable evaluation of agentic systems requires unbiased estimates with valid uncertainty, but standard practice navigates between costly human annotation and biased LLM-as-judge proxies. Prediction-powered inference (PPI) combines both into debiased estimates with valid confidence intervals, yet it…
- Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?
Ojas Nimase, Jiate Li, Yue Zhao, Yushun Dong · 1 juin 2026
Graph Machine Learning as a Service (GMLaaS) platforms increasingly implement explainability interfaces to meet regulatory transparency requirements. However, this transparency creates exploitable vulnerabilities for model extraction attacks. We present the first model extraction attack specifically…
- Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach
Chanwoo Park, Ziyang Chen, Asuman Ozdaglar, Kaiqing Zhang · 1 juin 2026
Large language models (LLMs) are increasingly deployed as "agents" for decision-making (DM) in interactive and dynamic environments. Yet, since they were not originally designed for DM, recent studies show that LLMs can struggle even in basic online DM problems, failing to achieve low regret or an e…
- The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Cassandra Goldberg, Chaehyeon Kim, Adam Stein, Eric Wong · 1 juin 2026
Concept vectors aim to enhance model interpretability by linking internal representations with human-understandable semantics, but their practical utility is often limited by noisy and inconsistent activations. In this work, we uncover the SuperActivator Mechanism: a transformer dynamic that amplifi…
- Self-Trained Verification for Training- and Test-Time Self-Improvement
Chen Henry Wu, Aditi Raghunathan · 29 mai 2026
Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall…
- RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
Haoxiang Jiang, Zihan Dong, Tianci Liu, Wanying Wang, Ran Xu, Tony Yu, Linjun Zhang, Haoyu Wang · 29 mai 2026
Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. Rubric-based methods address this by decomposing evaluation into explicit criteria, but existing approaches typically depend on frontier LLMs and suffe…
- ExDBSCAN: Explaining DBSCAN with Counterfactual Reasoning -- Additional Material
Pernille Matthews, Lena Krieger, Tommaso Amico, Artur Zimek, Thomas Seidl, Ira Assent · 29 mai 2026
Clustering is an unsupervised technique for grouping data points by similarity. While explainability methods exist for supervised machine learning, they are not directly applicable to clustering, making it challenging to understand cluster assignments. This interpretability gap is particularly evide…
- Anytime-Valid Federated Conformal RAG for LLM Swarms
Prasanjit Dubey, Xiaoming Huo · 29 mai 2026
Federated Conformal RAG (FC-RAG) provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizon. We extend it to anytime-valid sequential coverage: validity at every stopping time, preserved under predictable adaptive control (recalibration, pe…
- The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
Shu Wan, Abhinav Gorantla, Huan Liu, K. Sel\c{c}uk Candan · 29 mai 2026
Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant. Once the boundary is observed, the target is conditionally independent of the rest of the table. This is a tempting object for tabular prediction…
- Honest Lying: Understanding Memory Confabulation in Reflexive Agents
Prakhar Dixit, Sadia Kamal, Tim Oates · 29 mai 2026
Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures.We show that this assumption can fail systematically: across ALFWorld and HumanEval, agents store confident but incorrect interpretations of the task and co…
- CB-SLICE: Concept-Based Interpretable Error Slice Discovery
Yael Konforti, Mateo Espinosa Zarlenga, Elaf Almahmoud, Mateja Jamnik · 29 mai 2026
Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Identifying these groups and the root causes of their failures is critical for model debugging and bias mitigation. However, existing error Slice Discov…
