Cross-Cohort Robustness of Machine-Learning Gene-Expression Biomarkers for Lung Adenocarcinoma and Squamous Cell Carcinoma

Authors

  • Areen Jain Gretchen Whitney High School Author

Keywords:

lung adenocarcinoma, lung squamous cell carcinoma, gene expression, Elastic Net, biomarker stability, external validation, cross-cohort robustness

Abstract

Gene-expression classifiers can separate lung adenocarcinoma (LUAD) from lung squamous cell carcinoma (LUSC), but high apparent accuracy does not guarantee reproducible biomarkers. We developed a locked, cross-cohort analysis that combined nested cross-validation, Elastic-Net feature selection, resampling-based stability analysis, compact-panel evaluation, and a processing-to-platform shift ladder. The discovery cohort was GSE37745 (106 LUAD and 66 LUSC samples); the prespecified primary external cohort was GSE50081 (127 LUAD and 42 LUSC samples). A frozen Elastic-Net model achieved an internal out-of-fold ROC-AUC of 0.985 and an external ROC-AUC of 0.909 (DeLong 95% CI, 0.849–0.970). The external AUC degradation of 0.076 exceeded the prespecified 0.05 tolerance, showing that strong discrimination did not fully transfer without loss. Of 83 model-selected genes, 58 were selected in at least 80% of 100 stratified subsamples. Selection frequency was positively associated with external effect size (Spearman ρ= 0.385, p<0.001), and 10- and 20-gene panels retained 98.0% and 98.5% of the full-model external AUC, respectively. Performance remained high across several microarray cohorts but fell to near chance in TCGA RNA-seq data (AUC 0.511), emphasizing platform shift as a major boundary of generalization. These results support stability and external validation as essential complements to internal accuracy in transcriptomic biomarker studies.

Downloads

Published

2026-10-01