Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2.226 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady · 23. Juli 2026
Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores must reflect stable propert…
- What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education
Kai Yao · 23. Juli 2026
Generative AI is changing a basic premise of educational assessment: that submitted work can reliably evidence the human capacities a credential claims to certify. The challenge is not simply whether students use AI, but what remains inferable about learning when some cognitive work has been delegat…
- Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking
Ralf Raumanns, Theresa Elstner, Louis Ferger-Andrews, Louise M. Carlsen, Martin Potthast, Gerard Schouten, Josien P. W. Pluim, Veronika Cheplygina · 23. Juli 2026
Machine learning courses often use pre-labeled datasets, hiding the subjectivity of human annotation. This creates students with an overly trusting view of AI data and models, undervaluing interpretive diversity. We investigated whether manual data annotation tasks teach students about subjective la…
- The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks
Antonio Di Cecco · 23. Juli 2026
Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. We introduce the quadrilateral loss, a differentiable penalty that treats additivity as a measurable behavior instead: a second-order mixed difference on pairs of…
- Self-Explaining Reinforcement Learning for Mobile Network Resource Allocation
Konrad Nowosadko, Franco Ruggeri, Ahmad Terra · 23. Juli 2026
Deep reinforcement learning (DRL) methods, though powerful, often lack transparency, which limits their adoption in critical domains. We apply Self-Explaining Neural Networks (SENNs) to RL by parametrizing the policy of a PPO agent with a SENN, producing intrinsic local explanations, and propose a m…
- Using LLMs for Explainable, Data-Driven Insight Generation from Time Series
Ria Mundhra, Gustavo Sato dos Santos, Michael Benedikt · 22. Juli 2026
Time series forecasts are widely used in decision-critical domains, where they are rarely consumed without accompanying explanations. Producing such explanations is usually a manual and costly process, and attempts to automate it using large language models often suffer from hallucination when appli…
- Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime
Kushal Chakrabarti · 22. Juli 2026
As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account - more data, retrieval, or scale - misses an auto-regressive risk residual that scale sharpens: the model commits to a low-probability token, conditions on it a…
- Weak-to-Strong Learning in Decision Making
Jingwei Ji, Renyuan Xu · 22. Juli 2026
Many operational decisions rely on predictive models that estimate uncertain outcomes conditional on observable contexts. Training such models, however, often faces a fundamental data asymmetry: labeled outcomes are scarce or costly to obtain, while contextual covariates are abundant. Motivated by t…
- Spectral Higher-Order Neural Networks Have Sharp Expressivity Bounds
Gianluca Peri, Diego Febbe, Duccio Fanelli · 22. Juli 2026
Neural hypergraphs are a natural generalization of neural networks, the reference models in modern machine learning. Yet, their deployment has proven demanding: the number of weighted hyperedges required leads to an intractable parameter explosion. However, a novel parametrization that leverages spe…
- Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning
Wentao Zhang, Haoyu Zhang, Xinke Jiang, Yuxuan Cheng, Yuhan Pan, Miao Li, Zhipeng Qiao, Tao Feng, Zhen Tao, Dengji Zhao · 22. Juli 2026
Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to a…
- Calibrated Selective Fact-Checking via Evidence Chain Evaluation
Dekun Yang · 22. Juli 2026
Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts even when supporting evidence is weak, sparse, or internally inconsistent. We address this issue through Evidence Chain Eva…
- ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series
Annemarie Jutte, Faizan Ahmed, Jeroen Linssen, Maurice van Keulen · 22. Juli 2026
This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains such as healthcare and predictive maintenance, artificial intelligence models can increase efficiency and safety. Explainability is key to ensure these models r…
- A Geometric Perspective on Stabilizing Value Conflict Resolution
Saket Reddy, Andy Liu · 21. Juli 2026
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To address this challenge, we investigate how chain-of-thought (CoT) reasoning can help improve performance in this domain. Ge…
- RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
Mihir Shriniwas Arya · 21. Juli 2026
Large language models and LLM-based agents are widely used as personal chat assistants, enterprise copilots, and autonomous workflow agents. In all these applications, memory (the ability to retain, access, and reason over information accumulated over long contexts and multiple interactions) plays a…
- Human Grounded Evaluation of Large Language Models for Optical Network Automation
Kiarash Rezaei, Omran Ayoub, Paolo Monti, Carlos Natalino · 21. Juli 2026
Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable sc…
- A Method for Learning Value Systems in Generative AI
Andr\'es Holgado-S\'anchez, Holger Billhardt, Sascha Ossowski · 21. Juli 2026
Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems in order to align their decisions with ours. As such representations are difficult to elicit, value learning seeks to infer them by observing human behaviour. T…
- Expected Free Energy as Belief-Dependent Utility for rho-POMDPs
Patrick Cooper, Alvaro Velasquez · 21. Juli 2026
An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard POMDPs value information only through its eventual effect on reward. The $\rho$-POMDP framework instead rewards uncertainty reduction directly, through a belief-de…
- FST.ai 2.5: Explainable and Uncertainty-Aware AI for Olympic and Para-Taekwondo Decision Support, Athlete Digital Twins, and Federation-Scale Analytics
Keivan Shariatmadar, Ahmad Osman, Ramin Rey · 21. Juli 2026
The rapid digitalisation of elite sport has created new opportunities for integrating artificial intelligence (AI), performance analytics, and decision-support systems into athlete development and competition management. However, existing solutions remain fragmented, typically addressing isolated ta…
- Can Interpretation Predict Behavior on Unseen Data?
Victoria R. Li, Jenny Kaufmann, Tian Qin, Martin Wattenberg, David Alvarez-Melis, Naomi Saphra · 21. Juli 2026
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) behavior. We train hundreds of Tr…
- Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies
Luyu Qiu, Jianing Li, Hwanhee Kim, Xiaoyong Wei, Yueyuan Zheng, Janet Hsiao, Lei Chen · 21. Juli 2026
Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks. However, their subpar performance on seemingly elementary problems, such as basic arithmetic, raises concerns about model reliability, safety, and ethical…
- What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach
Afiq Abdillah Effiezal Aswadi, Haotong Ma, Susan Wei · 21. Juli 2026
A Bayes-filtered transformer (BFT) is a transformer trained on sequences that are generated in two steps: first a latent task is drawn from a prior, then observations are drawn conditional on that task. Trained under autoregressive log loss, the BFT's next-token prediction, in the idealized limit, i…
- Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
Damien Teney, Liangze Jiang, Hemanth Saratchandran, Simon Lucey · 21. Juli 2026
Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for pushing AI beyond merely scaling current designs. *Method.* We present a method to optimize a transformer arc…
- The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations
Amadeo Tunyi · 21. Juli 2026
Explainability methods for time series models predominantly produce flat attribution scores: they quantify the direct influence of a feature at a timestamp by a scalar. We prove that the dominant failure mode of such methods is not the scalar format itself but a fundamental computational mismatch: e…
- Counterfactual Shapley Credit Assignment
Mingxuan Li, Kaizhan-Lee, Elias Bareinboim · 21. Juli 2026
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently fail to attribute properly between an agent's policy (ski…
- Value-Monotonicity Matters: A Concordance Loss for Deep Survival Prediction
Meixu Chen, Kai Wang, Jing Wang · 21. Juli 2026
Deep survival models are evaluated almost exclusively by the concordance index (C-index), yet they are commonly trained using likelihood objectives such as the Cox partial likelihood, discrete-time negative log-likelihood, and DeepHit likelihood. This mismatch is usually considered acceptable becaus…
