Life Sciences › Biochemistry, Genetics and Molecular Biology › Molecular Biology
Biomedical Text Mining and Ontologies
284 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States46% · 56 papers
- China26% · 32 papers
- United Kingdom12% · 14 papers
- Germany9.1% · 11 papers
- Japan7.4% · 9 papers
- Canada7.4% · 9 papers
- Switzerland5.8% · 7 papers
- India5% · 6 papers
Across 121 papers on this subject with at least one lab located. 36 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes
Maryam Shahbaz Ali, Laura B. Strachan, Caitlin Sherman, Mark Kovler, Eleanor Mackey, Syed Muhammad Anwar · 2 October 2026
Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whe…
- ASCRIBE: Atomic and Significance-Based Reasoning for Thai Clinical SOAP Note Generation
Tarm Kalavantavanich, Teerawut Ponarchar, Pattaramanee Arsomngern, Jenta Wonglertsakul, Watcharakorn Chuthong, Chiraphat Boonnag, Knot Pipatsrisawat, Titipat Achakulvisut · 2 October 2026
Automatic SOAP note generation can ease the documentation burden on physicians, but existing reasoning methods often omit clinically important information and generate unsupported content. Progress in Thai is further hindered by the lack of publicly available datasets. We propose ASCRIBE, a physicia…
- Can large language models unlock discrete data in ophthalmic diagnostic reports?
Umair A. Zaidi, An-Lun Wu, Wei-Chun Lin, Thomas S. Hwang, Michelle R. Hribar · 2 October 2026
Objective: To assess the accuracy and efficiency of a large language model (LLM) using two prompt strategies to extract structured data from ophthalmic diagnostic PDF reports. Methods: Twenty deidentified reports across four types (Visual Field, OCT Glaucoma Overview, OCT retinal nerve fiber layer S…
- UniBuc at SemEval-2024 Task 2: Tailored Prompting with Solar for Clinical NLI
Marius Micluta-Campeanu, Claudiu Creanga, Ana-Maria Bucur, Ana Sabina Uban, Liviu P. Dinu · 2 October 2026
This paper describes the approach of the UniBuc team in tackling the SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials. We used SOLAR Instruct, without any fine-tuning, while focusing on input manipulation and tailored prompting. By customizing prompts for individua…
- Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?
Laura van Weesep, Riccardo Tedoldi, Jens Sj\"olund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, Juan Viguera Diez · 2 October 2026
The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this wo…
- LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction
Nouha Hayouni, Sheeba Samuel, Alsayed Algergawy · 2 October 2026
Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for onto…
- On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence
Vinay Kumar Chaganti · 2 October 2026
Named-entity recognition (NER) is increasingly wanted on-device (no API, low latency, data kept local). The practitioner's question is not the leaderboard but which model is deployable, how to evaluate it without human annotation, and whether its confidence can be trusted. We answer these jointly. W…
- An ontology for cross-sectoral crisis management: core and public health modules
Aldo Gangemi, Rita T. Sousa, Luigi Asprino, Giorgia Lodi, Andrea G. Nuzzolese, Valentina Presutti, Johannes Gysen, Diana F. Sousa, Luigi Spagnolo · 2 October 2026
This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental cr…
- Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Martin Krallinger, Miguel Rodr\'iguez-Ortega, Eduard Rodriguez-L\'opez, Natalia Loukachevitch, Igor Rozhkov, Elena Tutubalina, Dimitris Dimitriadis, Vasiliki Patsiou, Grigorios Tsoumakas, George Giannakoulas, Alexandra Bekiaridou, Athanasios Samaras, Giorgio Maria Di Nunzio, Nicola Ferro, Stefano Marchesin, Marco Martinelli, Gianmaria Silvello, Georgios Paliouras · 1 October 2026
This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic…
- When Scientific Contradictions Are Lost in Translation
Tal Zeevi, Trey W. Jensen, Maxwell Strome · 1 October 2026
Two scientific findings can disagree without contradicting each other. Determining whether they conflict requires knowing whether they describe comparable measurements. We study how language models behave at this decision point. In a controlled task, we generate an unsatisfiable XOR constraint syste…
- Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study
Bibek Bhandari, Kshitij Lingthep · 30 September 2026
Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5…
- The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond
Andreas Maier, Monica Hinrichs-Mayer, Franziska Weber, Niklas Lackner, Matthias May, Bernhard Kainz, Siming Bayer · 30 September 2026
Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymi…
- OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen, MingYu Lu, Huan-Yu Hsu, Yu Gu, Yue Guo, Sheng Wang, Wei Qiu, Hanwen Xu · 30 September 2026
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across…
- Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction
Azrin Sultana · 30 September 2026
The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced genes lacking characterized functions. Traditional homology-based methods are often costly and strongly dependent on high sequence similarity. This study prese…
- Follow the Entities: A Corpus Map for Agentic Search
Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam · 30 September 2026
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively se…
- TRACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs
Jizheng Lai, Yingyun Li, Ying Qin, Haiyang Qian · 30 September 2026
Large language models are increasingly used in oncology applications, but their predictions are often weakly grounded in explicit medical structure. We present TRACE, a deployable tree-relational enhancement framework for oncology LLMs. TRACE separates expensive offline structure learning from light…
- OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang · 30 September 2026
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluatio…
- ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering
Yuyan Chen · 30 September 2026
In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive …
- MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning
Lang Cao, Binghang Lu, Yuhao Shen, Yue Guo · 29 September 2026
Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any si…
- What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents
Zhaowei Han, Xiang Zhang, Lingxiao Guan, Danqi Hu, Kai Liu, Kevin Chang, Jie Liu · 29 September 2026
Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We…
- Extraction of clinical findings from mammography and breast ultrasound reports: a comparison between specialists and Artificial Intelligence
Lorenzo Farias, Hanna Reckziegel, Daniela Duarte da Silva Bagatini, Daniel Schulz, Gabriela de Andrade Monteiro, Let\'icia Zanatta, Ana Laura Brill Thum, Priscila Schmidt Lora, D\'ebora Oliveira da Silva, Ana Paula Wernz da Cunha M\"uller, Cristiane Drebes Pedron · 29 September 2026
Breast cancer is the leading cause of cancer-related death among women in Brazil, and the time between the request and the release of mammography reports directly influences adherence to screening, making the agility in processing these reports a critical factor for early diagnosis. In this context,…
- Toward AI-Assisted Poultry Coccidiosis Diagnosis: Evaluating Gemini and BiomedParse on Eimeria Microscopy Images
Ali Alsalama, Ahmed Kubba, Manar Abu Talib · 29 September 2026
Coccidiosis caused by Eimeria parasites is a major economic burden in poultry production, and effective control depends on accurate species-level diagnosis. This study evaluates whether a general-purpose multimodal large language model can support such diagnosis. Google Gemini was assessed on 4,225 …
- DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery
Yulong Li, Rong Xia, Yuxuan Zhang, Jianxu Chen, Xiwei Liu, Haochen Xue, Maosheng Li, Yuhang Liu, Yibo Yuan, Yutong Xie, Chong Li, Jionglong Su, Hagai Rossman, Eran Segal, Imran Razzak · 29 September 2026
We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, fro…
- Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten, P. Ellen Grant, Yangming Ou · 29 September 2026
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL…
- DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?
Nan Huang, Mario Tapia-Pacheco, Kun Zhou, Yiming Huang, Kevin Jos\'e Barrientos D\'iaz, Tiffany Amariuta, Jingbo Shang · 29 September 2026
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical ta…
