Life Sciences › Biochemistry, Genetics and Molecular Biology › Molecular Biology
Biomedical Text Mining and Ontologies
284 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos46 % · 56 artículos
- China26 % · 32 artículos
- Reino Unido12 % · 14 artículos
- Alemania9,1 % · 11 artículos
- Japón7,4 % · 9 artículos
- Canadá7,4 % · 9 artículos
- Suiza5,8 % · 7 artículos
- India5 % · 6 artículos
Sobre 121 artículos de este tema con al menos un laboratorio localizado. 36 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study
Yun Wang, Gad Shaulsky, Toma\v{z} Curk, Bla\v{z} Zupan · 5 de octubre de 2026
Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introdu…
- Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories
Joseph Chee Chang, Michael D'Arcy, Amy X. Zhang, Pao Siangliulue, Sangho Suh, Aakanksha Naik, Jena D. Hwang, Javier Ramos Benitez, Stella Wroblewski, Matt Latzke, Michael Cuoco, Ruben Lozano-Aguilera, Kris Ganjam, Joel Chan, Doug Downey, Peter Jansen, Kyle J. Travaglini, Daniel S. Weld · 5 de octubre de 2026
A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their prefer…
- A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition
Songtao Li, Yijia Zhang, Shidi Zhang, Jianyuan Yuan, Fengyu Zhang, Hongfei Lin · 5 de octubre de 2026
Large language models (LLMs) have shown promising potential for biomedical named entity recognition (BioNER) through instruction following and in-context learning. However, existing LLM-based BioNER methods still face two key limitations. First, retrieved demonstrations and external biomedical knowl…
- Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method
Songtao Li, Yijia Zhang, Jianyuan Yuan, Shidi Zhang, Fengyu Zhang, Hongfei Lin · 5 de octubre de 2026
Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annot…
- From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
Pritam Deka · 5 de octubre de 2026
Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transf…
- Toward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells, Organoids, and Biobots
Nam H. Le, Douglas Blackiston, Michael Levin, Josh Bongard · 5 de octubre de 2026
Artificial intelligence increasingly serves as a natural-language interface to complex technical systems, letting people accomplish sophisticated tasks by describing what they want rather than specifying how to do it. Extending this interface to living systems is harder: unlike code or images, a bio…
- Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models
Manan Roy Choudhury, Suparno Roy Chowdhury, Swastik Sahoo, Muhammad Ali Khan, Kaneez Zahra Rubab Khakwani, Mohamad Bassam Sonbol, Irbaz Bin Riaz, Vivek Gupta · 5 de octubre de 2026
Systematic reviews condense clinical trials into evidence tables, yet clinicians can interrogate these tables only through database queries, and many questions concern attributes that the table does not record, such as a drug's target class or a harmonised endpoint. Here we introduce FD-SCoPE, a lan…
- Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes
Maryam Shahbaz Ali, Laura B. Strachan, Caitlin Sherman, Mark Kovler, Eleanor Mackey, Syed Muhammad Anwar · 2 de octubre de 2026
Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whe…
- ASCRIBE: Atomic and Significance-Based Reasoning for Thai Clinical SOAP Note Generation
Tarm Kalavantavanich, Teerawut Ponarchar, Pattaramanee Arsomngern, Jenta Wonglertsakul, Watcharakorn Chuthong, Chiraphat Boonnag, Knot Pipatsrisawat, Titipat Achakulvisut · 2 de octubre de 2026
Automatic SOAP note generation can ease the documentation burden on physicians, but existing reasoning methods often omit clinically important information and generate unsupported content. Progress in Thai is further hindered by the lack of publicly available datasets. We propose ASCRIBE, a physicia…
- Can large language models unlock discrete data in ophthalmic diagnostic reports?
Umair A. Zaidi, An-Lun Wu, Wei-Chun Lin, Thomas S. Hwang, Michelle R. Hribar · 2 de octubre de 2026
Objective: To assess the accuracy and efficiency of a large language model (LLM) using two prompt strategies to extract structured data from ophthalmic diagnostic PDF reports. Methods: Twenty deidentified reports across four types (Visual Field, OCT Glaucoma Overview, OCT retinal nerve fiber layer S…
- UniBuc at SemEval-2024 Task 2: Tailored Prompting with Solar for Clinical NLI
Marius Micluta-Campeanu, Claudiu Creanga, Ana-Maria Bucur, Ana Sabina Uban, Liviu P. Dinu · 2 de octubre de 2026
This paper describes the approach of the UniBuc team in tackling the SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials. We used SOLAR Instruct, without any fine-tuning, while focusing on input manipulation and tailored prompting. By customizing prompts for individua…
- Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?
Laura van Weesep, Riccardo Tedoldi, Jens Sj\"olund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, Juan Viguera Diez · 2 de octubre de 2026
The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this wo…
- LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction
Nouha Hayouni, Sheeba Samuel, Alsayed Algergawy · 2 de octubre de 2026
Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for onto…
- On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence
Vinay Kumar Chaganti · 2 de octubre de 2026
Named-entity recognition (NER) is increasingly wanted on-device (no API, low latency, data kept local). The practitioner's question is not the leaderboard but which model is deployable, how to evaluate it without human annotation, and whether its confidence can be trusted. We answer these jointly. W…
- An ontology for cross-sectoral crisis management: core and public health modules
Aldo Gangemi, Rita T. Sousa, Luigi Asprino, Giorgia Lodi, Andrea G. Nuzzolese, Valentina Presutti, Johannes Gysen, Diana F. Sousa, Luigi Spagnolo · 2 de octubre de 2026
This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental cr…
- Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Martin Krallinger, Miguel Rodr\'iguez-Ortega, Eduard Rodriguez-L\'opez, Natalia Loukachevitch, Igor Rozhkov, Elena Tutubalina, Dimitris Dimitriadis, Vasiliki Patsiou, Grigorios Tsoumakas, George Giannakoulas, Alexandra Bekiaridou, Athanasios Samaras, Giorgio Maria Di Nunzio, Nicola Ferro, Stefano Marchesin, Marco Martinelli, Gianmaria Silvello, Georgios Paliouras · 1 de octubre de 2026
This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic…
- When Scientific Contradictions Are Lost in Translation
Tal Zeevi, Trey W. Jensen, Maxwell Strome · 1 de octubre de 2026
Two scientific findings can disagree without contradicting each other. Determining whether they conflict requires knowing whether they describe comparable measurements. We study how language models behave at this decision point. In a controlled task, we generate an unsatisfiable XOR constraint syste…
- Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study
Bibek Bhandari, Kshitij Lingthep · 30 de septiembre de 2026
Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5…
- The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond
Andreas Maier, Monica Hinrichs-Mayer, Franziska Weber, Niklas Lackner, Matthias May, Bernhard Kainz, Siming Bayer · 30 de septiembre de 2026
Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymi…
- OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen, MingYu Lu, Huan-Yu Hsu, Yu Gu, Yue Guo, Sheng Wang, Wei Qiu, Hanwen Xu · 30 de septiembre de 2026
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across…
- Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction
Azrin Sultana · 30 de septiembre de 2026
The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced genes lacking characterized functions. Traditional homology-based methods are often costly and strongly dependent on high sequence similarity. This study prese…
- Follow the Entities: A Corpus Map for Agentic Search
Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam · 30 de septiembre de 2026
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively se…
- TRACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs
Jizheng Lai, Yingyun Li, Ying Qin, Haiyang Qian · 30 de septiembre de 2026
Large language models are increasingly used in oncology applications, but their predictions are often weakly grounded in explicit medical structure. We present TRACE, a deployable tree-relational enhancement framework for oncology LLMs. TRACE separates expensive offline structure learning from light…
- OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang · 30 de septiembre de 2026
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluatio…
- ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering
Yuyan Chen · 30 de septiembre de 2026
In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive …
