Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2.319 indexierte Paper
Erklärbare künstliche Intelligenz zielt darauf ab, die Entscheidungen und Mechanismen von KI-Modellen verständlich zu machen, insbesondere durch die Identifizierung der Elemente, die ihre Vorhersagen beeinflussen. Aktuelle Arbeiten untersuchen Methoden wie Shapley values, Aktivierungskarten oder kausale Zerlegungen, um Daten oder internen Strukturen von Systemen - sei es neuronalen Netzen, Transformern oder Multi-Agenten-Systemen - eine präzise Rolle zuzuweisen. Dieser Ansatz hinterfragt auch die Robustheit der Erklärungen, ihre Konsistenz gegenüber analytischen Variationen oder Störungen sowie die Art und Weise, wie Konzepte der Kausalität oder interner Mechanismen disziplinübergreifend interpretiert werden.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten40 % · 658 Artikel
- China22 % · 352 Artikel
- Deutschland11 % · 173 Artikel
- Vereinigtes Königreich8,6 % · 141 Artikel
- Kanada5,4 % · 88 Artikel
- Indien5,2 % · 85 Artikel
- Frankreich4,4 % · 72 Artikel
- Italien3,5 % · 57 Artikel
Über 1.633 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 84 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift
Ruben Dario Florez-Zela · 5. Oktober 2026
Explainability is increasingly required for perception models in intelligent vehicles, yet whether explanations remain faithful under driving domain shift is still poorly understood. This work audits post-hoc explanations of a fixed YOLOv8s pedestrian detector across PIE and JAAD using ROI-based D-D…
- Beyond Trained Models: Compiling GNNs for a Sound Explainer Benchmark
Steve Azzolin, Francesco Paolo Nerini, Stefano Teso, Francesco Bonchi, Bruno Lepri, Andr\'e Panisson, Andrea Passerini · 5. Oktober 2026
Explainers for Graph Neural Networks (GNNs) are commonly evaluated by their plausibility, i.e., how well their explanations recover a predefined ground truth, such as a motif planted in the data. This protocol implicitly assumes that a GNN trained on such data relies on the intended motif. Although …
- MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
Yuyang Cheng, Raghav Kaushik Ravi, Srivarshinee Sridhar, Sriparna Saha, Akash Ghosh, Chirag Agarwal · 5. Oktober 2026
Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands ex…
- How the Audit Rule Shapes Faithful Factor Explanations in LLMs
Taolin Zhang, Hanyu Wang, Jiuheng Wan, Tingyuan Hu, Chengyu Wang · 2. Oktober 2026
Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this l…
- Compliant AI Infrastructure for Regulated Finance: A tiered multi-agent framework with DLT audit trails for financial operations in DACH
Walter Kurz, Reinhard Magg · 2. Oktober 2026
We present a compliance-first architecture for AI in regulated finance that treats regulation as an orientation layer rather than a deterministic ruleset. A matrix of regulatory intent and exposure provides a compact classification handle, which a governed policy compiler then maps into concrete pro…
- MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees
Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey, Yan Zhang · 2. Oktober 2026
Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstabl…
- Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention
Guy Amit · 2. Oktober 2026
Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressi…
- A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
Javier Diaz Esteban-Herreros, David Mu\~noz-Valero, Raquel Mart\'inez-Espa\~na, Jose M. Juarez, Juan Moreno-Garcia · 2. Oktober 2026
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and p…
- Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers
Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen · 2. Oktober 2026
Dynamic architectures leave a per-example routing trace beside each prediction, and diffuse routing is easy to read as a sign that the prediction is unreliable. We audit that reading for routing entropy in Attention-Residual (AR) variants of Swin-Tiny and DeiT-Small, trained from scratch on CIFAR-10…
- Out-of-Distribution Detection using Counterfactual Distance
Maria Stoica, Francesco Leofante, Alessio Lomuscio · 1. Oktober 2026
Accurate and explainable out-of-distribution (OOD) detection is required to use machine learning systems safely. Previous work has shown that feature distance to decision boundaries can be used to identify OOD data effectively. In this paper, we build on this intuition and propose a post-hoc OOD det…
- Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis
Xia Hu, Brian Potetz, Chun-Ta Lu, Huanfen Yao, Leonidas Guibas, Zhicheng Wang, Howard Zhou, Pengfei Xing, Andrew Gallagher · 1. Oktober 2026
End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure m…
- Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations
Hanwei Zhang, Tianma Hu, Gaojie Jin, Xu Cheng, Ronghui Mu · 1. Oktober 2026
Concept Bottleneck Models (CBMs) are designed to provide interpretable intermediate representations, yet how such bottlenecks affect robustness remains unclear, with existing studies reporting mixed and sometimes contradictory findings. We argue that these discrepancies arise from conflating differe…
- Routing Probes Can Improve Without New Information: An Exact-Null Audit of Uncertainty Beyond Model Outputs
Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen · 1. Oktober 2026
Routing signals of modern vision transformers -- expert gates, attention-residual weights and halting scores -- often improve probes that predict whether the model is correct, and the improvement is commonly read as evidence that routing carries information about errors beyond the model's outputs. W…
- When the Wrong Key Wins: Understanding and Detecting Hallucinations in LLMs
Xuhan Tong, Haoyue Bai, Dawei Zhou, Naichen Shi, Jiawei Zhang · 30. September 2026
Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a latent-key view of inference, where answer selection depends on competition among associations acquired during pretraining. We show that model predictions…
- Frontier Learning: Training LLM Reasoners at the Edge of Capability
Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi · 30. September 2026
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as …
- Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts
Chenxiao Fan, Chongming Gao, Gangyi Zhang, Leyang Shen, Yaxin Gong, Jiamin Wang, Jiakai Wang, Dong Wang, Yang Liu, Fuli Feng, Xiangnan He · 30. September 2026
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model s…
- Steering Language Model Goals with Value Transplant
Pengcheng Jiang, Fabien Roger · 30. September 2026
Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their goals through a "value axis." We study …
- From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features
Dewen Liu, Zixuan Li, Jonathan Pan, Zhao Wu, Zijun Yao, Juanzi Li, Xiaozhi Wang · 30. September 2026
Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while …
- RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
Jingyi He, Nier Wu, Shuang Liu, Xin Wang, Mengnan Du, Xia Hu · 30. September 2026
Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their sc…
- Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models
Jonathan Drechsel, Steffen Herbold · 30. September 2026
Model-internal features can be studied through both their ability to identify a specified concept and their causal effect when manipulated, e.g., through steering or weight editing. A prominent approach to feature learning is Sparse Autoencoders (SAEs), which learn broad feature dictionaries whose r…
- COLLATOR: Compositional Multi-Agent Orchestration with Counterfactual Reinforcement Learning
Xudong Chen, Yixin Liu, Hua Wei, Kaize Ding · 30. September 2026
Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend critically on orchestration design. Across different tasks, role design, capacity assignment, and dependency construction jointly affect both solution quality and…
- The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
Gon\c{c}alo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman · 30. September 2026
Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying…
- Selecting The Most Informative Tokens in Natural Language Autoencoders
Federico Torrielli, Gianluca Barmina, Andrea Blasi N\'u\~nez, Amon Rapp, Luigi Di Caro, Peter Schneider-Kamp, Lukas Galke Poech · 30. September 2026
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt inject…
- Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability
Zhenting Huang, Junnan Liu, Qianren Mao, Zhixing Tan, Bo Jiang · 30. September 2026
Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms…
- Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training
Adam Elimadi · 30. September 2026
Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an o…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
- Natural Language Processing Techniques1.595 Papiere / 12 Monate+69 %
