Life Sciences › Biochemistry, Genetics and Molecular Biology › Molecular Biology
Gene expression and cancer classification
17 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Neueste Paper
- Hierarchical Sparse Bayesian Multitask Learning for Disease Prediction in Pooled Microbiome Studies
Haonan Zhu, Andre R. Goncalves, Camilo Valdes, Hiranmayi Ranganathan, Boya Zhang, Jose Manuel Mart\'i, Car Reen Kok, Monica K. Borucki, Nisha J. Mulakken, James B. Thissen, Crystal Jaing, Alfred Hero, Nicholas A. Be · 23. September 2026
This paper proposes a hierarchical Bayesian multitask learning model that is applicable to the general multi-task binary classification learning problem where the model assumes a shared sparsity structure across different tasks. We derive a computationally efficient inference algorithm based on vari…
- Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
Alexandros Tzanakakis, Aris Karatzikos, Ilias Georgakopoulos-Soares · 17. September 2026
A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foundation models and whether structural geometry determines functional importance remains unknown. We analyzed high-gain rows in gate…
- Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining
Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz · 31. August 2026
As a precursor to high-dimensional biomedical data modeling, reliable feature selection can reduce computational expense, improve modeling performance, and yield simpler, more interpretable models. However, most filter-based feature selection methods struggle to detect feature interactions, while wr…
- ARISE: An adaptive residual-informed stability ensemble for feature selection in small-sample biomedical omics
Zardad Khan, Amjad Ali, Naz Gul, Sheema Gul, Saeed Aldahmani · 18. August 2026
Objective: Small-sample molecular classification requires feature selectors that identify predictive, stable, and nonredundant subsets for binary and multiclass outcomes. We propose ARISE (Adaptive Residual-Informed Stability Ensemble), which integrates complementary relevance signals, class-balance…
- Nonlinear multi-study sparse factor analysis
Gemma E. Moran, Anandi Krishnan · 12. August 2026
High-dimensional data often exhibit variation that can be captured by lower-dimensional factors. For high-dimensional data from multiple studies, one goal is to understand which underlying factors are common to all studies, and which factors are study-specific. As a particular example, we consider p…
- Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study
Sepideh Saran, Mahsa Ghanbari, Uwe Ohler · 12. August 2026
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work p…
- Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays
Xiaohua Douglas Zhang · 11. August 2026
High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-learning performance metrics, such as sensitivity, sp…
- The Cost of Binarizing Survival Outcomes in Clinical Prognostic Modeling
Shashank Yadav, David M. Routman, Andrew Y. K. Foong · 6. August 2026
Survival analysis is an established framework for analyzing time-to-event data, yet many clinical machine learning studies still binarize the outcome before model training. This practice excludes censored patients, collapses temporal information into a single threshold, and can affect which features…
- Evaluation and Prognostic Validation of Deep Regression Models for WSI-Based Gene-Expression Prediction
Fredrik K. Gustafsson, Constance Boissin, Johan Vallon-Christersson, Mattias Rantalainen · 24. Juli 2026
Gene-expression profiling is widely used in research and central to many areas of precision oncology, but remains costly and not universally accessible. Recent advances in computational pathology enable prediction of transcriptomic profiles directly from hematoxylin and eosin (H&E)-stained whole-sli…
- Cross-Cluster Weighted Forests
Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani · 17. Juli 2026
Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies. We propose the 'Cross-Cluster Weighted Forest' (CCWF), an ensembling approach that explicitly leverages heterogeneity in th…
- When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection
Jesus S. Aguilar-Ruiz · 1. Juli 2026
Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret. Variables are first ranked by a relevance score, and a subset is then obtained by retaining the top-ranked variables. Although the first stage has been extensively studied, the s…
- ERICA: Quantifying Replicability of Cluster Analysis
Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani · 2. Juni 2026
Despite being ubiquitous in science, clustering remains a technique whose results are not quantitatively scrutinized via a framework. We present an analysis called evaluating replicability via iterative clustering assignments (ERICA) that is applied to a dataset to determine whether clusters are ide…
- Innovative Silicosis and Pneumonia Classification: Leveraging Graph Transformer Post-hoc Modeling and Ensemble Techniques
Bao Q. Bui, Tien T. T. Nguyen, Duy M. Le, Cong Tran, Cuong Pham · 27. Mai 2026
This paper presents a comprehensive study on the classification and detection of Silicosis-related lung inflammation. Our main contributions include 1) the creation of a newly curated chest X-ray (CXR) image dataset named SVBCX that is tailored to the nuances of lung inflammation caused by distinct …
- Querying structural and functional niches on spatial transcriptomics data
Mo Chen, Minsheng Hao, Xinquan Liu, Lin Deng, Peng Liu, Chen Li, Dongfang Wang, Kui Hua, Liang Guo, Xuegong Zhang, Lei Wei · 26. Mai 2026
Cells in multicellular organisms coordinate to form structural and functional niches. With spatial transcriptomics (ST) enabling gene expression profiling in spatial contexts, it has been revealed that spatial niches serve as cohesive and recurrent units in physiological and pathological processes. …
- RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction
Yaxuan Song, Jianan Fan, Tianyi Wang, Qiuyue Hu, Hang Chang, Heng Huang, Weidong Cai · 13. Mai 2026
Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular architecture and functional programs defining pathological states, whereas RNA sequencing (RNA-seq) provides genome-wide transcriptional profiles at subst…
- BioBlobs: Unsupervised Discovery of Functional Substructures for Protein Function Prediction
Xin Wang, Kaiwen Shi, Carlos Oliver · 13. Mai 2026
Protein function is driven by cohesive substructures, such as catalytic triads, binding pockets, and structural motifs, that occupy only a small fraction of a protein's residues. Yet existing pipelines built on protein encoders do not model proteins at the substructure level, leaving the central bio…
- Feature Dimensionality Outweighs Model Complexity in Breast Cancer Subtype Classification Using TCGA-BRCA Gene Expression Data
Meena Al Hasani · 8. Mai 2026
Accurate classification of breast cancer subtypes from gene expression data is critical for diagnosis and treatment selection. However, such datasets are characterized by high dimensionality and limited sample size, posing challenges for machine learning models. In this study, we evaluate the impa…
- Disease Is a Spectral Perturbation
John D. Mayfield, Matthew S. Rosen · 6. Mai 2026
We propose a novel method of understanding disease transformation from a healthy baseline with biomarker-level explainability. By modeling the biomarker covariance matrices of healthy controls and disease states, the perturbation can be individually characterized to accomplish mechanistic explanatio…
- Biconvex Biclustering
Sam Rosen, Eric C. Chi, Jason Xu · 7. April 2026
This article proposes a biconvex modification to convex biclustering in order to improve its performance in high-dimensional settings. In contrast to heuristics that discard a subset of noisy features a priori, our method jointly learns and accordingly weighs informative features while discovering b…
- KGroups: A Versatile Univariate Max-Relevance Min-Redundancy Feature Selection Algorithm for High-dimensional Biological Data
Malick Ebiele, Malika Bendechache, Rob Brennan · 31. März 2026
This paper proposes a new univariate filter feature selection (FFS) algorithm called KGroups. The majority of work in the literature focuses on investigating the relevance or redundancy estimations of feature selection (FS) methods. This has shown promising results and a real improvement of FFS meth…
- Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization
Yupei Zhang, Xiaofei Wang, Anran Liu, Lequan Yu, Chao Li · 3. März 2026
Histopathology remains the gold standard for cancer diagnosis and prognosis. With the advent of transcriptome profiling, multi-modal learning combining transcriptomics with histology offers more comprehensive information. However, existing multi-modal approaches are challenged by intrinsic multi-mod…
- Supervised Graph Contrastive Learning for Gene Regulatory Networks
Sho Oshima, Yuji Okamoto, Taisei Tosaki, Ryosuke Kojima · 20. Februar 2026
Graph Contrastive Learning (GCL) is a powerful self-supervised learning framework that performs data augmentation through graph perturbations, with growing applications in the analysis of biological networks such as Gene Regulatory Networks (GRNs). The artificial perturbations commonly used in GCL, …
- MAFS: Multi-head Attention Feature Selection for High-Dimensional Data via Deep Fusion of Filter Methods
Xiaoyan Sun, Qingyu Meng, Yalu Wen · 7. Januar 2026
Feature selection is essential for high-dimensional biomedical data, enabling stronger predictive performance, reduced computational cost, and improved interpretability in precision medicine applications. Existing approaches face notable challenges. Filter methods are highly scalable but cannot capt…
- MFAI: A Scalable Bayesian Matrix Factorization Approach to Leveraging Auxiliary Information
Zhiwei Wang, Fa Zhang, Cong Zheng, Xianghong Hu, Mingxuan Cai, Can Yang · 6. Januar 2026
In various practical situations, matrix factorization methods suffer from poor data quality, such as high data sparsity and low signal-to-noise ratio (SNR). Here, we consider a matrix factorization problem by utilizing auxiliary information, which is massively available in real-world applications, t…
- Sparse Convex Biclustering
Jiakun Jiang, Dewei Xiang, Chenliang Gu, Wei Liu, Binhuan Wang · 6. Januar 2026
Biclustering is an essential unsupervised machine learning technique for simultaneously clustering rows and columns of a data matrix, with widespread applications in genomics, transcriptomics, and other high-dimensional omics data. Despite its importance, existing biclustering methods struggle to me…
