Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 219 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Dewi Gould, Francis Rhys Ward, Anders Cairns Woodruff, Rauno Arike, Josh Hills, Alex Serrano, Ida Caspary, Jason Ross Brown, Jo J. Jiao, Patrick Leask, Twm Stone, Ram Potham, Ionut Gabriel Stan, Harry Mayne, Simeon Hellsten, Shubhorup Biswas, Ariana Azarbal, William L. Anderson, Elle Najt, Ryan Greenblatt, Julian Stastny · 8 juin 2026
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning internally, without explicit thinking tokens, this would undermine such oversight. We measure how well frontier models reason …
- An Abstract Architecture for Explainable Autonomy in Hazardous Environments
Matt Luckcuck, Hazel M Taylor, Marie Farrell · 8 juin 2026
Autonomous robotic systems are being proposed for use in hazardous environments, often to reduce the risks to human workers. In the immediate future, it is likely that human workers will continue to use and direct these autonomous robots, much like other computerised tools but with more sophisticate…
- Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers
Tang Li, Yanlin Chen, Mengmeng Ma, Xi Peng · 8 juin 2026
Despite high accuracy, Vision Transformer (ViT) predictions can be driven by spurious cues, raising the need to understand their inner workings before safe deployment. Sparse autoencoders (SAEs) provide a promising lens for decomposing model representations into human-interpretable concepts, yet ada…
- How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures
Tanvi Thoria, Kiana Jafari, Marc R. Schlichting, Mykel J. Kochenderfer · 8 juin 2026
Failures in language model reasoning emerge through distinct processes that leave identifiable signatures in the reasoning trace. We characterize these failures using token-level uncertainty signals, finding they arise through two empirically distinguishable processes. The first is committed failure…
- Perplexity Can Miss SAE Feature Damage Under Quantization
Evan Duan · 8 juin 2026
Quantization is a standard path to deploying large language models, and quantized models are typically judged acceptable when perplexity or downstream accuracy remains close to the full-precision original. But behavioral parity need not imply feature fidelity: the sparse-autoencoder (SAE) features u…
- Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces
Debjyoti Saha Roy, Byron C. Wallace, Javed A. Aslam · 8 juin 2026
Modern reasoning models offer surprisingly strong zero-shot performance on challenging multi-label tasks that require selecting a small set of relevant options from hundreds of thousands to millions of candidate labels. We investigate how they achieve this mechanistically. We characterize reasoning …
- Automatic, Debiased, and Invariant Counterfactual Generation under General Interventions
Raphael C Kim, Jingsen Zhu, Ramin Zabih, Michele Santacatterina · 8 juin 2026
Generative models for counterfactual outcomes have great potential to support decision-making under complex interventions, but existing approaches are limited by unstable estimation, poor generalization across environments, and bias from nuisance model misspecification. We introduce ADIGen, a framew…
- I-SAFE: Wasserstein Coherence Metrics for Structural Auditing of Scientific AI Models
Barbara Tarantino, Gennaro Auricchio, Paolo Giudici · 8 juin 2026
Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scientifically meaningful behavior. This interpretation is fragile, as models may exploit shortcut features, dataset-specific regularities, or distribution…
- OPTIMUS-Prime: Minimal and Sufficient Concept Explanations for Deep Vision Models
Arthur Hoarau, Chenrui Zhu, Vu Linh Nguyen · 8 juin 2026
The growing demand for transparency in automated decision-making has propelled eXplainable Artificial Intelligence (XAI) to the forefront of machine learning research. In computer vision, however, existing explanation methods often prioritize end-user accessibility at the expense of formal guarantee…
- Sparsely gated tiny linear experts
Simon Schug · 8 juin 2026
Sparsity allows scaling model parameters without proportionally increasing computational cost. While mixture of experts (MoE) models are made increasingly sparse, individual experts typically remain large and dense. Here, we demonstrate that further increasing sparsity by shrinking each expert to co…
- A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders
Chenhao Zhang, Chris Lin, Su-In Lee · 8 juin 2026
We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpretability of neural networks by learning sparse feature representations, a principled definition of ''concept'' and ''learn…
- Self-evolving LLM agents with in-distribution Optimization
Yudi Zhang, Meng Fang, Zhenfang Chen, Mykola Pechenizkiy · 8 juin 2026
Large Language Models (LLMs) have recently emerged as powerful controllers for interactive agents in complex environments, yet training them to perform reliable long-horizon decision making remains a fundamental challenge. A key difficulty lies in credit assignment: agents often receive delayed rewa…
- Bias in Filter Feature Selection Evaluation: A Meta-Analysis of Datasets, Baselines, and Experimental Design Choices
Malick Ebiele, Malika Bendechache, Rob Brennan · 8 juin 2026
Background: Since 1990 many feature selection methods have been proposed across heterogeneous applications. To validate the usefulness of a new method, it needs to be compared against at least one baseline method from the existing literature on a feature selection task using at least one dataset. Re…
- Beyond Post-hoc Explanation: Toward Glassbox AI via Probabilistic Mediation
Manuele Leonelli · 8 juin 2026
Large language models are rapidly becoming infrastructural components in high-stakes institutional settings, including public administration, legal reasoning, and healthcare, where opacity is not merely inconvenient but institutionally and legally untenable. Existing approaches to explainability are…
- StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents
Haojie Hao, Longkun Hao, Yihang Lou, Yan Bai, Zhenyang Li, Zhichao Yang, Dongshuo Huang, Hongyu Lin, Lanqing Hong, Jiakai Wang, Xianglong Liu · 8 juin 2026
Reinforcement Learning (RL) has become a promising approach for improving GUI Agents in long-horizon, stochastic digital environments, but trajectory-level success feedback is too sparse to provide reliable credit assignment for intermediate exploration steps. To mitigate this issue, recent studies …
- Metamorphic Testing with the Rashomon Set: Explanation Faithfulness in Machine Learning
Helge Spieker, J{\o}rn Eirik Betten, Arnaud Gotlieb · 5 juin 2026
Multiple machine learning models can achieve near-equivalent predictive performance on the same task, yet provide divergent feature-based explanations. This is called the Rashomon effect of (explainable) machine learning, and it raises the question of which explanations, if any, are trustworthy. We …
- Beyond Soft Masks: Hard-Perturbation Mixup Explainer for Robust GNN Explainability
Jialiang Yin, Zheng Zhao, Linsey Pang, Bo Dong, Bin Shi, Jiaxing Zhang · 5 juin 2026
Graph Neural Networks (GNNs) have demonstrated remarkable performance across a range of applications involving graph-structured data, particularly in high-stakes domains. However, the opaque nature of their decision-making processes limits their trustworthiness and broader adoption. Existing post-ho…
- Towards AI epidemiology: a measurement standardisation framework for prospective risk detection
Kit Tempest-Walters · 5 juin 2026
This paper proposes a measurement standardisation framework that compresses expert-AI interactions into structured, comparable fields for prospective risk detection in deployed AI systems, without access to model internals. The main aim of this concept paper is to define the scope of the framework, …
- A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
Ranjan Mishra, Jakob Schoeffer · 5 juin 2026
Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusively on point predictions as AI advice. However, set-valued AI advice (e.g., discrete sets or continuous intervals) is increasingly being used to communicate uncer…
- From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents
Patrick Wilhelm, Odej Kao · 5 juin 2026
Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal model state and environment context. We study reward-hacking monitors in ReAct-style agents acting in Gameable ALFWorld and WebShop. Agents are instrume…
- The Self-Correction Illusion: LLMs Correct Others but Not Themselves
Kuan-Yen Chen, Fang-Yi Su, Jung-Hsien Chiang · 5 juin 2026
Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a capability deficit or a role-label artifact: does an agent's willingness …
- CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives
Ayoung Lee, Ryan Sungmo Kwon, Peter Railton, Lu Wang · 5 juin 2026
Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been limited to everyday scenarios. To close this gap, we introduce CLASH (Character perspective-based LLM Assessments in Situations with High-stakes), a metic…
- Fix the Mind, Not the Move: Interpretable AI Assistance via Knowledge-Gap Localization
Ayano Hiranaka, Ya-Chuan Hsu, Stefanos Nikolaidis, Erdem B{\i}y{\i}k, Daniel Seita · 5 juin 2026
AI assistants in human-AI collaboration often correct suboptimal human actions through behavioral feedback (e.g., alerts or steering-wheel nudges in assistive driving). Such interventions can mitigate immediate errors, but long-term improvement requires addressing the underlying misconceptions that …
- Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents
Zeyu Gan, Huayi Tang, Yong Liu · 5 juin 2026
As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With the rapid expansion of available skills, enabling personal agents to learn and adapt to implicit user preferences becomes…
- Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral Patterns
Olasimbo Ayodeji Arigbabu · 5 juin 2026
AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of agent behavior: whether an agent explores too much, repeats itself too rigidly, uses tools effectively, reduces uncertainty over time, or remains robust …
