Life Sciences › Biochemistry, Genetics and Molecular Biology › Molecular Biology
Gene expression and cancer classification
17 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Latest papers
- $l_{1-2}$ GLasso: $L_{1-2}$ Regularized Multi-task Graphical Lasso for Joint Estimation of eQTL Mapping and Gene Network
Wei Miao, Lan Yao · 29 September 2026
A critical problem in genetics is to discover how gene expression is regulated within cells. Two major tasks of regulatory association learning are : (i) identifying SNP-gene relationships, known as eQTL mapping, and (ii) determining gene-gene relationships, known as gene network estimation. To shar…
- Hierarchical Sparse Bayesian Multitask Learning for Disease Prediction in Pooled Microbiome Studies
Haonan Zhu, Andre R. Goncalves, Camilo Valdes, Hiranmayi Ranganathan, Boya Zhang, Jose Manuel Mart\'i, Car Reen Kok, Monica K. Borucki, Nisha J. Mulakken, James B. Thissen, Crystal Jaing, Alfred Hero, Nicholas A. Be · 23 September 2026
This paper proposes a hierarchical Bayesian multitask learning model that is applicable to the general multi-task binary classification learning problem where the model assumes a shared sparsity structure across different tasks. We derive a computationally efficient inference algorithm based on vari…
- Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
Alexandros Tzanakakis, Aris Karatzikos, Ilias Georgakopoulos-Soares · 17 September 2026
A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foundation models and whether structural geometry determines functional importance remains unknown. We analyzed high-gain rows in gate…
- Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining
Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz · 31 August 2026
As a precursor to high-dimensional biomedical data modeling, reliable feature selection can reduce computational expense, improve modeling performance, and yield simpler, more interpretable models. However, most filter-based feature selection methods struggle to detect feature interactions, while wr…
- ARISE: An adaptive residual-informed stability ensemble for feature selection in small-sample biomedical omics
Zardad Khan, Amjad Ali, Naz Gul, Sheema Gul, Saeed Aldahmani · 18 August 2026
Objective: Small-sample molecular classification requires feature selectors that identify predictive, stable, and nonredundant subsets for binary and multiclass outcomes. We propose ARISE (Adaptive Residual-Informed Stability Ensemble), which integrates complementary relevance signals, class-balance…
- Nonlinear multi-study sparse factor analysis
Gemma E. Moran, Anandi Krishnan · 12 August 2026
High-dimensional data often exhibit variation that can be captured by lower-dimensional factors. For high-dimensional data from multiple studies, one goal is to understand which underlying factors are common to all studies, and which factors are study-specific. As a particular example, we consider p…
- Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study
Sepideh Saran, Mahsa Ghanbari, Uwe Ohler · 12 August 2026
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work p…
- Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays
Xiaohua Douglas Zhang · 11 August 2026
High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-learning performance metrics, such as sensitivity, sp…
- The Cost of Binarizing Survival Outcomes in Clinical Prognostic Modeling
Shashank Yadav, David M. Routman, Andrew Y. K. Foong · 6 August 2026
Survival analysis is an established framework for analyzing time-to-event data, yet many clinical machine learning studies still binarize the outcome before model training. This practice excludes censored patients, collapses temporal information into a single threshold, and can affect which features…
- Evaluation and Prognostic Validation of Deep Regression Models for WSI-Based Gene-Expression Prediction
Fredrik K. Gustafsson, Constance Boissin, Johan Vallon-Christersson, Mattias Rantalainen · 24 July 2026
Gene-expression profiling is widely used in research and central to many areas of precision oncology, but remains costly and not universally accessible. Recent advances in computational pathology enable prediction of transcriptomic profiles directly from hematoxylin and eosin (H&E)-stained whole-sli…
- Cross-Cluster Weighted Forests
Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani · 17 July 2026
Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies. We propose the 'Cross-Cluster Weighted Forest' (CCWF), an ensembling approach that explicitly leverages heterogeneity in th…
- When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection
Jesus S. Aguilar-Ruiz · 1 July 2026
Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret. Variables are first ranked by a relevance score, and a subset is then obtained by retaining the top-ranked variables. Although the first stage has been extensively studied, the s…
- ERICA: Quantifying Replicability of Cluster Analysis
Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani · 2 June 2026
Despite being ubiquitous in science, clustering remains a technique whose results are not quantitatively scrutinized via a framework. We present an analysis called evaluating replicability via iterative clustering assignments (ERICA) that is applied to a dataset to determine whether clusters are ide…
- Innovative Silicosis and Pneumonia Classification: Leveraging Graph Transformer Post-hoc Modeling and Ensemble Techniques
Bao Q. Bui, Tien T. T. Nguyen, Duy M. Le, Cong Tran, Cuong Pham · 27 May 2026
This paper presents a comprehensive study on the classification and detection of Silicosis-related lung inflammation. Our main contributions include 1) the creation of a newly curated chest X-ray (CXR) image dataset named SVBCX that is tailored to the nuances of lung inflammation caused by distinct …
- Querying structural and functional niches on spatial transcriptomics data
Mo Chen, Minsheng Hao, Xinquan Liu, Lin Deng, Peng Liu, Chen Li, Dongfang Wang, Kui Hua, Liang Guo, Xuegong Zhang, Lei Wei · 26 May 2026
Cells in multicellular organisms coordinate to form structural and functional niches. With spatial transcriptomics (ST) enabling gene expression profiling in spatial contexts, it has been revealed that spatial niches serve as cohesive and recurrent units in physiological and pathological processes. …
- RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction
Yaxuan Song, Jianan Fan, Tianyi Wang, Qiuyue Hu, Hang Chang, Heng Huang, Weidong Cai · 13 May 2026
Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular architecture and functional programs defining pathological states, whereas RNA sequencing (RNA-seq) provides genome-wide transcriptional profiles at subst…
- BioBlobs: Unsupervised Discovery of Functional Substructures for Protein Function Prediction
Xin Wang, Kaiwen Shi, Carlos Oliver · 13 May 2026
Protein function is driven by cohesive substructures, such as catalytic triads, binding pockets, and structural motifs, that occupy only a small fraction of a protein's residues. Yet existing pipelines built on protein encoders do not model proteins at the substructure level, leaving the central bio…
- Feature Dimensionality Outweighs Model Complexity in Breast Cancer Subtype Classification Using TCGA-BRCA Gene Expression Data
Meena Al Hasani · 8 May 2026
Accurate classification of breast cancer subtypes from gene expression data is critical for diagnosis and treatment selection. However, such datasets are characterized by high dimensionality and limited sample size, posing challenges for machine learning models. In this study, we evaluate the impa…
- Disease Is a Spectral Perturbation
John D. Mayfield, Matthew S. Rosen · 6 May 2026
We propose a novel method of understanding disease transformation from a healthy baseline with biomarker-level explainability. By modeling the biomarker covariance matrices of healthy controls and disease states, the perturbation can be individually characterized to accomplish mechanistic explanatio…
- Biconvex Biclustering
Sam Rosen, Eric C. Chi, Jason Xu · 7 April 2026
This article proposes a biconvex modification to convex biclustering in order to improve its performance in high-dimensional settings. In contrast to heuristics that discard a subset of noisy features a priori, our method jointly learns and accordingly weighs informative features while discovering b…
- KGroups: A Versatile Univariate Max-Relevance Min-Redundancy Feature Selection Algorithm for High-dimensional Biological Data
Malick Ebiele, Malika Bendechache, Rob Brennan · 31 March 2026
This paper proposes a new univariate filter feature selection (FFS) algorithm called KGroups. The majority of work in the literature focuses on investigating the relevance or redundancy estimations of feature selection (FS) methods. This has shown promising results and a real improvement of FFS meth…
- Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization
Yupei Zhang, Xiaofei Wang, Anran Liu, Lequan Yu, Chao Li · 3 March 2026
Histopathology remains the gold standard for cancer diagnosis and prognosis. With the advent of transcriptome profiling, multi-modal learning combining transcriptomics with histology offers more comprehensive information. However, existing multi-modal approaches are challenged by intrinsic multi-mod…
- Supervised Graph Contrastive Learning for Gene Regulatory Networks
Sho Oshima, Yuji Okamoto, Taisei Tosaki, Ryosuke Kojima · 20 February 2026
Graph Contrastive Learning (GCL) is a powerful self-supervised learning framework that performs data augmentation through graph perturbations, with growing applications in the analysis of biological networks such as Gene Regulatory Networks (GRNs). The artificial perturbations commonly used in GCL, …
- MAFS: Multi-head Attention Feature Selection for High-Dimensional Data via Deep Fusion of Filter Methods
Xiaoyan Sun, Qingyu Meng, Yalu Wen · 7 January 2026
Feature selection is essential for high-dimensional biomedical data, enabling stronger predictive performance, reduced computational cost, and improved interpretability in precision medicine applications. Existing approaches face notable challenges. Filter methods are highly scalable but cannot capt…
- MFAI: A Scalable Bayesian Matrix Factorization Approach to Leveraging Auxiliary Information
Zhiwei Wang, Fa Zhang, Cong Zheng, Xianghong Hu, Mingxuan Cai, Can Yang · 6 January 2026
In various practical situations, matrix factorization methods suffer from poor data quality, such as high data sparsity and low signal-to-noise ratio (SNR). Here, we consider a matrix factorization problem by utilizing auxiliary information, which is massively available in real-world applications, t…
