An Optimised Computational Framework for Cross-Cohort Metagenomic Classification Using P-Value Statistical Filtering
Abstract
Metagenomic data analysis faces significant challenges due to high dimensionality, sparsity, and geographical heterogeneity, which limit the generalizability of diagnostic models. This study proposes a robust computational framework that integrates p-value-based statistical filtering with ensemble machine learning to identify stable microbial signatures for colorectal cancer (CRC). By ranking features via T-tests, we identified an optimal subset of 170 features, significantly reducing dimensionality from the original 2000 features while enhancing predictive performance. Results demonstrate high system efficiency: the Gradient Boosting Classifier (GBC) achieved an Area Under the Receiver Operating Characteristic Curve (AUC) of 0.88 on the Zeller cohort, compared to 0.80 using the full feature set. Extensive cross-cohort evaluations across four global datasets (Germany/France, Austria, China, and USA) identified the Zeller dataset as a “Gold Standard” for feature extraction, yielding biomarkers that are highly transferable across diverse populations. Furthermore, ensemble models (GBC and Random Forest) exhibited superior resilience to distribution shifts compared to SVM, which suffered from performance collapse in cross-dataset scenarios. This framework provides a computationally lean and scalable solution for large-scale metagenomic studies, emphasising the importance of accounting for population-specific variability in robust system design.
© 2026 Hien Thanh Thi Nguyen, Hai Thanh Nguyen, published by Riga Technical University
This work is licensed under the Creative Commons Attribution 4.0 License.