Cross-Cohort Robustness of Machine-Learning Gene-Expression Biomarkers for Lung Adenocarcinoma and Squamous Cell Carcinoma
Keywords:
lung adenocarcinoma, lung squamous cell carcinoma, gene expression, Elastic Net, biomarker stability, external validation, cross-cohort robustnessAbstract
Gene-expression classifiers can separate lung adenocarcinoma (LUAD) from lung squamous cell carcinoma (LUSC), but high apparent accuracy does not guarantee reproducible biomarkers. We developed a locked, cross-cohort analysis that combined nested cross-validation, Elastic-Net feature selection, resampling-based stability analysis, compact-panel evaluation, and a processing-to-platform shift ladder. The discovery cohort was GSE37745 (106 LUAD and 66 LUSC samples); the prespecified primary external cohort was GSE50081 (127 LUAD and 42 LUSC samples). A frozen Elastic-Net model achieved an internal out-of-fold ROC-AUC of 0.985 and an external ROC-AUC of 0.909 (DeLong 95% CI, 0.849–0.970). The external AUC degradation of 0.076 exceeded the prespecified 0.05 tolerance, showing that strong discrimination did not fully transfer without loss. Of 83 model-selected genes, 58 were selected in at least 80% of 100 stratified subsamples. Selection frequency was positively associated with external effect size (Spearman ρ= 0.385, p<0.001), and 10- and 20-gene panels retained 98.0% and 98.5% of the full-model external AUC, respectively. Performance remained high across several microarray cohorts but fell to near chance in TCGA RNA-seq data (AUC 0.511), emphasizing platform shift as a major boundary of generalization. These results support stability and external validation as essential complements to internal accuracy in transcriptomic biomarker studies.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Areen Jain (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.