Health Sciences › Medicine › Health Informatics
Artificial Intelligence in Healthcare and Education
1075 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos47 % · 330 artículos
- China21 % · 150 artículos
- Reino Unido11 % · 75 artículos
- Canadá7,7 % · 54 artículos
- Alemania6,8 % · 48 artículos
- India5,5 % · 39 artículos
- Australia4,3 % · 30 artículos
- Francia3,3 % · 23 artículos
Sobre 705 artículos de este tema con al menos un laboratorio localizado. 74 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- In Dialogue with Intelligence: Toward Insightful Co-Augmentation
Eleni Vasilaki · 2 de octubre de 2026
Dialogue with a large language model can lead a person to insight: a sudden change in how they understand a problem. This perspective asks how model activity relates to insight as a dialogue unfolds. I propose that part of the intelligence expressed in dialogue arises from two interacting recurrence…
- Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage
Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang · 2 de octubre de 2026
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinica…
- OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou · 2 de octubre de 2026
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review bec…
- A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance
HyungJun Kim, Taehan Lee, Soojin Cheon · 2 de octubre de 2026
Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, exe…
- CALLIOPE: A Source-Grounded Oral Assessment System and Synthetic Readiness Evaluation
Nizam Kadir · 2 de octubre de 2026
Oral assessment with generative AI requires more than a conversational interface: educators must connect a spoken response to its source material, scoring criteria, model outputs and subsequent human judgement. This technical report presents CALLIOPE, a source-grounded oral assessment system integra…
- ContextAdapt: Evaluating Contextual Adaptation and Value Alignment in LLMs
Olivia Macmillan-Scott, Mirco Musolesi · 1 de octubre de 2026
Values such as honesty, autonomy, and confidentiality are often regarded as general principles underpinning AI alignment. However, what it means to act in accordance with these values can depend on the context in which a decision is made. In this paper, we ask whether large language models (LLMs) ap…
- What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents
Shuyang Zhang (The Hong Kong Polytechnic University), Jianshuo Chang (The Hong Kong Polytechnic University) · 1 de octubre de 2026
Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. Thi…
- Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi · 1 de octubre de 2026
As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlyin…
- Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach
Sandra Woolley, Tim Collins, Khalid Khattak, Illia Chernomorets, Ariane Arevalo, Chris Richardson · 1 de octubre de 2026
This paper examines human-AI interactions (HAIIs) in clinical trials and presents a multidimensional categorisation framework that classifies interactions according to AI tasks, human-AI relationships, interaction configurations and interacting human groups. We define HAII, examine existing taxonomi…
- Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
Anastasiia Ivanova, Natalia Fedorova, Ekaterina Artemova · 1 de octubre de 2026
The rapid development of AI-driven tools, particularly large language models (LLMs), is reshaping professional writing. Still, key aspects of their adoption such as language support, ethics, and long-term impact on writers' voice and creativity remain underexplored. In this work, we carried out a qu…
- SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing
Dhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh · 30 de septiembre de 2026
SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, and workplace chat, sampling each model on…
- KinyaMed: Seeds, Not Rows -- What a Corpus Requirement Written in the Wrong Unit Fails to Constrain
Marius Bayizere · 30 de septiembre de 2026
Triage decides who is seen first. Building an urgency classifier for patient-voice Kinyarwanda, we found our specification could be met without producing anything it was meant to secure. We report that, and the instruments that detect it, instead of a classifier. Designed for the four languages a Rw…
- Single-turn emergency psychiatric triage across 15 frontier AI chatbots
Veith Weilnhammer, Lennart Luettgau, Christopher Summerfield, Raymond Dolan, Elise Wilkinson, Virginia Corno, Viknesh Sounderajah, Matthew M Nour · 30 de septiembre de 2026
People increasingly turn to general-purpose AI chatbots for advice about emotional and mental health problems, but the ability of these systems to recognize and appropriately triage psychiatric emergencies remains under-characterized. We evaluated psychiatric triage performance in 15 frontier AI c…
- A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo · 30 de septiembre de 2026
Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, He…
- Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models
Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon · 30 de septiembre de 2026
Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, part…
- Accessible, but Not Adopted: Increasing LLM Adoption among First-generation, Low-income (FGLI) College Students beyond Expanding Access
Hyungsik Kim · 30 de septiembre de 2026
Large language models (LLMs) are increasingly positioned as a force to empower underserved communities, and significant efforts are being made to expand access. Yet, access alone does not equate to meaningful adoption. First, even if a system is accessible, it won't be adopted if users are not willi…
- Large Language Models Exhibit Human-Like Bayesian Hypocrisy
Nykko Vitali, Mahzarin R. Banaji · 30 de septiembre de 2026
Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observe…
- Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
Nicol\'as Vera Z\'u\~niga · 29 de septiembre de 2026
Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LL…
- Tool Mediation Alters Refusal Mechanisms in Large Language Models
Abel Rodr\'iguez, Giuseppe Garofalo, Lieven Desmet, Vera Rimmer · 29 de septiembre de 2026
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms …
- Applying Language Models in medical Medicine: Recent Trends and Perspectives
Erik Aerts · 29 de septiembre de 2026
The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While …
- Jev in Medicine: A Benchmark Evaluation. Preliminary Results
Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho · 29 de septiembre de 2026
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaM…
- Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support
Bharath Kumar Bolla, Bharath Kumar Bolla, Vishnu Surya Reddy Nandi · 29 de septiembre de 2026
Clinical Large Language Models (LLMs) achieve strong medical-exam accuracy; however, a correct answer does not guarantee that the explanation names the concepts that actually drove the decision. We introduce three lightweight, directly interpretable metrics for this faithfulness gap: the Explanation…
- Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs
Chaitai Deb Purkayastha, Bharath Kumar Bolla, Vishnu Surya Reddy Nandi · 29 de septiembre de 2026
Clinical LLM evaluation often emphasizes answer accuracy; however, accuracy alone does not test counterfactual consistency or demographic robustness. We evaluated six LLMs on 150 MedQA USMLE questions using two automated perturbation tests to assess their performance. The counterfactual validity (CF…
- A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases
Linkai Li, Changgeng Mo, Hanlin Yu, Congxi Lu, Shangqiguo Wang, Matthew B Fitzgerald, Shan X Wang · 29 de septiembre de 2026
Audiology consultation requires structured history-taking, audiometric interpretation and patient-centred communication, yet real-world case material is scarce. We present a bilingual AI audiologist pairing a general-purpose large language model with rubric-guided playbook induction, multimodal audi…
- HIPAA-Compliant AI Deployment Patterns in Clinical Settings: Privacy-Preserving Techniques and Governance Controls
Vinod Dhiman · 29 de septiembre de 2026
Artificial intelligence (AI) is moving from research prototypes into clinical workflows, yet the deployment of AI systems that process protected health information (PHI) remains constrained by the U.S. Health Insurance Portability and Accountability Act (HIPAA) and by the absence of shared engineeri…
