Health Sciences › Medicine › Health Informatics
Artificial Intelligence in Healthcare and Education
1.075 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten47 % · 330 Artikel
- China21 % · 150 Artikel
- Vereinigtes Königreich11 % · 75 Artikel
- Kanada7,7 % · 54 Artikel
- Deutschland6,8 % · 48 Artikel
- Indien5,5 % · 39 Artikel
- Australien4,3 % · 30 Artikel
- Frankreich3,3 % · 23 Artikel
Über 705 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 74 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models
Mehrdad Ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani, Sadra Hakim, Audrina Ebrahimi · 5. Oktober 2026
Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Per…
- Clinical Concept Centers in LLMs
Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler · 5. Oktober 2026
Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher f…
- In Dialogue with Intelligence: Toward Insightful Co-Augmentation
Eleni Vasilaki · 2. Oktober 2026
Dialogue with a large language model can lead a person to insight: a sudden change in how they understand a problem. This perspective asks how model activity relates to insight as a dialogue unfolds. I propose that part of the intelligence expressed in dialogue arises from two interacting recurrence…
- Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage
Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang · 2. Oktober 2026
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinica…
- OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou · 2. Oktober 2026
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review bec…
- A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance
HyungJun Kim, Taehan Lee, Soojin Cheon · 2. Oktober 2026
Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, exe…
- CALLIOPE: A Source-Grounded Oral Assessment System and Synthetic Readiness Evaluation
Nizam Kadir · 2. Oktober 2026
Oral assessment with generative AI requires more than a conversational interface: educators must connect a spoken response to its source material, scoring criteria, model outputs and subsequent human judgement. This technical report presents CALLIOPE, a source-grounded oral assessment system integra…
- ContextAdapt: Evaluating Contextual Adaptation and Value Alignment in LLMs
Olivia Macmillan-Scott, Mirco Musolesi · 1. Oktober 2026
Values such as honesty, autonomy, and confidentiality are often regarded as general principles underpinning AI alignment. However, what it means to act in accordance with these values can depend on the context in which a decision is made. In this paper, we ask whether large language models (LLMs) ap…
- What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents
Shuyang Zhang (The Hong Kong Polytechnic University), Jianshuo Chang (The Hong Kong Polytechnic University) · 1. Oktober 2026
Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. Thi…
- Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi · 1. Oktober 2026
As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlyin…
- Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach
Sandra Woolley, Tim Collins, Khalid Khattak, Illia Chernomorets, Ariane Arevalo, Chris Richardson · 1. Oktober 2026
This paper examines human-AI interactions (HAIIs) in clinical trials and presents a multidimensional categorisation framework that classifies interactions according to AI tasks, human-AI relationships, interaction configurations and interacting human groups. We define HAII, examine existing taxonomi…
- Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
Anastasiia Ivanova, Natalia Fedorova, Ekaterina Artemova · 1. Oktober 2026
The rapid development of AI-driven tools, particularly large language models (LLMs), is reshaping professional writing. Still, key aspects of their adoption such as language support, ethics, and long-term impact on writers' voice and creativity remain underexplored. In this work, we carried out a qu…
- SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing
Dhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh · 30. September 2026
SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, and workplace chat, sampling each model on…
- KinyaMed: Seeds, Not Rows -- What a Corpus Requirement Written in the Wrong Unit Fails to Constrain
Marius Bayizere · 30. September 2026
Triage decides who is seen first. Building an urgency classifier for patient-voice Kinyarwanda, we found our specification could be met without producing anything it was meant to secure. We report that, and the instruments that detect it, instead of a classifier. Designed for the four languages a Rw…
- Single-turn emergency psychiatric triage across 15 frontier AI chatbots
Veith Weilnhammer, Lennart Luettgau, Christopher Summerfield, Raymond Dolan, Elise Wilkinson, Virginia Corno, Viknesh Sounderajah, Matthew M Nour · 30. September 2026
People increasingly turn to general-purpose AI chatbots for advice about emotional and mental health problems, but the ability of these systems to recognize and appropriately triage psychiatric emergencies remains under-characterized. We evaluated psychiatric triage performance in 15 frontier AI c…
- A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo · 30. September 2026
Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, He…
- Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models
Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon · 30. September 2026
Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, part…
- Accessible, but Not Adopted: Increasing LLM Adoption among First-generation, Low-income (FGLI) College Students beyond Expanding Access
Hyungsik Kim · 30. September 2026
Large language models (LLMs) are increasingly positioned as a force to empower underserved communities, and significant efforts are being made to expand access. Yet, access alone does not equate to meaningful adoption. First, even if a system is accessible, it won't be adopted if users are not willi…
- Large Language Models Exhibit Human-Like Bayesian Hypocrisy
Nykko Vitali, Mahzarin R. Banaji · 30. September 2026
Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observe…
- Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
Nicol\'as Vera Z\'u\~niga · 29. September 2026
Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LL…
- Tool Mediation Alters Refusal Mechanisms in Large Language Models
Abel Rodr\'iguez, Giuseppe Garofalo, Lieven Desmet, Vera Rimmer · 29. September 2026
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms …
- Applying Language Models in medical Medicine: Recent Trends and Perspectives
Erik Aerts · 29. September 2026
The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While …
- Jev in Medicine: A Benchmark Evaluation. Preliminary Results
Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho · 29. September 2026
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaM…
- Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support
Bharath Kumar Bolla, Bharath Kumar Bolla, Vishnu Surya Reddy Nandi · 29. September 2026
Clinical Large Language Models (LLMs) achieve strong medical-exam accuracy; however, a correct answer does not guarantee that the explanation names the concepts that actually drove the decision. We introduce three lightweight, directly interpretable metrics for this faithfulness gap: the Explanation…
- Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs
Chaitai Deb Purkayastha, Bharath Kumar Bolla, Vishnu Surya Reddy Nandi · 29. September 2026
Clinical LLM evaluation often emphasizes answer accuracy; however, accuracy alone does not test counterfactual consistency or demographic robustness. We evaluated six LLMs on 150 MedQA USMLE questions using two automated perturbation tests to assess their performance. The counterfactual validity (CF…
