Physical Sciences › Computer Science › Artificial Intelligence
Explainable Artificial Intelligence (XAI)
2 219 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling
Heng Ping, Arijit Bhattacharjee, Peiyu Zhang, Shixuan Li, Wei Yang, Ali Jannesari, Nesreen Ahmed, Paul Bogdan · 24 juin 2026
Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents into layered reasoning pipelines. However, existing MoA variants fail to sustain gains as depth increases, exhibiting degradation, early plateauing, or saturation. We propose ReM-MoA, a memory-augm…
- You Don't Need to Run Every Eval
Yuchen Zeng, Dimitris Papailiopoulos · 24 juin 2026
A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release. But do we need to run every eval? We compile a public score matrix of 84 frontier models…
- Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir · 24 juin 2026
Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence. We present a human-grounded evaluation fra…
- Bayesian control for coding agents
Theodore Papamarkou, Vladislav Smirnov, Viktor Mazanov, Artem Vazhentsev, Preslav Nakov, Timothy Baldwin, Artem Shelmanov · 24 juin 2026
Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive verifiers. The tool-use decisions are typically governed by orchestrators that often use fixed rules and ignore uncertainty. We formulate orchestration as cost-sensitive sequential hypothesis testi…
- Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?
Ayan Antik Khan, Harsh Kohli, Yuekun Yao, Huan Sun, Ziyu Yao · 24 juin 2026
Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) agents can assist with this explanation problem once a…
- Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback
Andreas Chouliaras, Luke Connolly, Dimitris Chatzpoulos · 24 juin 2026
Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback. While both show promising results, no publicly av…
- When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
Hiroshi Okumura · 24 juin 2026
Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts. While prior benchmark studies have primarily evaluated LLMs' causal reasoning capabilities, a more fundamental epistemic dimension has been overlooked: Causal Caution, defined as the…
- Cycle-Consistent Neural Explanation of Formal Verification Certificates
Andoni Rodriguez, Alberto Pozanco, Daniel Borrajo · 24 juin 2026
Formal verification produces machine-checkable certificates that attest to the satisfaction or violation of temporal properties, yet these certificates remain opaque to non-specialist stakeholders. We propose a cycle-consistent neural architecture that generates faithful natural language explanation…
- BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei · 24 juin 2026
Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in individual tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of ho…
- Signed Evidence Flow: Conflict-Aware and Stability-Calibrated Data Analysis
Jeffery Opoku, David Banahene · 23 juin 2026
Modern data analysis usually gives a prediction without showing whether the evidence behind it is clear, conflicting, or stable. Two cases can have the same fitted confidence even when one has mostly agreeing evidence and the other has strong support and strong opposition. We propose Signed Evidence…
- Causal Discovery in the Era of Agents
Yujia Zheng, Vishal Verma, Mantej Gill, Haoyue Dai, Peter Spirtes, Kun Zhang · 23 juin 2026
Recent attempts to combine large language models (LLMs) with causal discovery ask models to infer pairwise directions, propose graph structures, or inject language-model outputs as priors and constraints. These approaches promise faster analysis, but they also obscure whether a causal evidence is su…
- Confidence-Uncertainty Boundary Calibration for Bayesian Deep Learning in Medical Image Analysis
Hua Xu, Juli\'an D. Arias-Londo\~no, Juan I. Godino-Llorente · 23 juin 2026
In critical decision support systems based on medical imaging, the reliability of AI-assisted decision-making is as relevant as predictive accuracy. Although deep learning models have demonstrated significant accuracy, they frequently suffer from miscalibration, manifested as overconfidence in erron…
- Where Is My Physics Wrong? Localized and Identifiable Discovery of Model Discrepancy
Yifan Wang · 23 juin 2026
Hybrid models combine trusted physics with data-driven correction, but a physical model is rarely wrong everywhere or in the same way. The key diagnostic question is local: where does the model fail, what missing mechanism explains the failure, and is the evidence statistically real? Existing sparse…
- Discovering Latent Groups for Robust Classification
Ankur Garg, Ulrich A\"ivodji, Samira Ebrahimi Kahou, Vincent Michalski · 23 juin 2026
Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately on underrepresented subgroups. Existing methods address this by adjusting network parameters, guided either by subgroup annotations or inferred pseudo-group labels. Yet at inference,…
- Null-Calibrated Conformal Selection via Target-Membership Scores
Seungjin Choi · 23 juin 2026
Conformal selection aims to identify test candidates whose unknown responses fall in a target region while controlling the false discovery rate. Existing methods often inherit prediction-oriented nonconformity scores, such as residual or clipped residual scores, from conformal prediction. We argue t…
- PrivacyAlign: Contextual Privacy Alignment for LLM Agents
Manveer Singh Tamber, Abhay Puri, Marc-Etienne Brunet, Perouz Taslakian, Jimmy Lin, Spandana Gella · 23 juin 2026
AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an important alignment problem for agents: every message, post, or tool call an agent makes is a contextual judgment about wha…
- On the Expressive Power of Weight Quantization in Large Language Models
Shao-Qun Zhang · 23 juin 2026
In recent years, weight quantization that encodes the learnable parameters of large language models in an $n$-bit format has garnered significant attention due to its potential for model compression and inference acceleration. Many practical techniques have been developed; however, the theoretical u…
- Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution
Soyeon Kim, Kyowoon Lee, Jaesik Choi · 23 juin 2026
Path-based attribution methods such as Integrated Gradients (IG) are widely adopted for their strong axiomatic properties and effectiveness in attributing model predictions to input features by integrating gradients along a path from a baseline to the input. However, the choice of the attribution pa…
- From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
Yifan Li, Shengbin Yue, Boyu Feng, Jinhu Qi, Bo Ke, Zixing Song, Hongru Wang, Zhongyu Wei, Irwin King · 23 juin 2026
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting self-awareness capability, the ability to discern whether a problem requires necessary external resources or can be solved…
- Measuring Behavior Portability in Large Language Models
Tianjia Dong, Nadav Kunievsky, James A. Evans · 23 juin 2026
Large language models are increasingly deployed as autonomous decision makers, yet the behavioral mapping they exhibit can vary substantially across decision environments that are payoff-equivalent by construction-environments that share identical payoff-relevant structure but differ in surface pres…
- Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents
Chubin Zhang, Zhenglin Wan, Xingrui Yu, Pengfei Zhou, Wangbo Zhao, Jingxuan Wu, Yaxin Zhou, Ivor Tsang · 23 juin 2026
Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would the agent be better off receiving no task evidence? We study this question with a controlled matched-loop comparison that…
- When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents
Aman Mehta · 23 juin 2026
Long-horizon LLM agents can fail quietly: they settle on one reading of the evidence early, then spend the rest of the run defending it. We call this premature commitment. Final-answer scoring misses the failure mode because it sees only the answer, not whether the process has already collapsed to a…
- HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs
Yucheng Wu, Jundong Xu, Mingzhen Ju, Yue Yu, Chenpeng Wang, Haoxuan Li, Liangming Pan · 23 juin 2026
Logical reasoning is essential for reliable AI, yet existing benchmarks are largely first-order-logic-centric, focusing on object-level deduction over fixed predicates. This misses many realistic scenarios where models must reason over rules, predicates, functions, constraints, and decision procedur…
- Human Decision-Making with AI Assistance under Correlated Features
Yanru Guan, Naveen Raman, Fei Fang · 23 juin 2026
Humans increasingly make decisions with AI assistance; for example, doctors may follow AI-recommended diagnostic tests and base their diagnoses on the results. A natural question is which tests should AI recommend to balance short-term decision quality and long-term human learning when different fea…
- Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers
Vatsal Ananthula, Adarsh Kumarappan · 23 juin 2026
Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reasoning. We propose verifier-coupled reasoning, a framework that inserts inline claims into reasoning traces and trains an auxiliary consistency head to…
