Biomarker discovery in inflammatory bowel diseases using network-based feature selection

Biomarker discovery in inflammatory bowel diseases using network-based feature selection
复制标题

DOI:
10.1371/journal.pone.0225382
复制
发表时间:
2019-11-22
期刊:
影响因子:
3.7
通讯作者:
EL-Manzalawy, Yasser
EL-Manzalawy, Yasser
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Abbas, Mostafa;Matta, John;EL-Manzalawy, Yasser

文献摘要

被引文献

相似文献

可靠地鉴定来自宏基因组学数据的炎症生物标志物是开发非侵入性,具有成本效益和快速临床测试的有希望的方向,用于早期诊断IBD。我们提出了一种基于网络的生物标志物发现(NBBD)的综合方法,该方法集成了网络分析,以优先考虑潜在的生物标志物和机器学习技术,以评估优先生物标志物的歧视能力。使用大量新的小儿IBD元基因组活检样本的数据集,我们比较了随机森林(RF)分类器的性能,该分类器(RF)分类器接受了使用代表性的传统功能选择方法针对NBBD框架进行选择的功能,该功能采用了针对NBBD框架,该方法使用五个不同的工具,用于推断网络来推断网络从元基因组学数据以及九种不同的方法来确定生物标志物以及结合最佳传统和NBBD的混合方法基于功能选择。我们还研究了IBD诊断预测模型的性能如何随着用于生物标志物识别的数据的大小而变化。我们的结果表明,(i)NBBD具有一些最先进的特征选择方法(包括随机森林特征重要性(RFFI)分数)具有竞争力; (ii)当可靠的数据样本数量较小时,NBBD在可靠地识别IBD生物标志物方面特别有效。
Reliable identification of Inflammatory biomarkers from metagenomics data is a promising direction for developing non-invasive, cost-effective, and rapid clinical tests for early diagnosis of IBD. We present an integrative approach to Network-Based Biomarker Discovery (NBBD) which integrates network analyses methods for prioritizing potential biomarkers and machine learning techniques for assessing the discriminative power of the prioritized biomarkers. Using a large dataset of new-onset pediatric IBD metagenomics biopsy samples, we compare the performance of Random Forest (RF) classifiers trained on features selected using a representative set of traditional feature selection methods against NBBD framework, configured using five different tools for inferring networks from metagenomics data, and nine different methods for prioritizing biomarkers as well as a hybrid approach combining best traditional and NBBD based feature selection. We also examine how the performance of the predictive models for IBD diagnosis varies as a function of the size of the data used for biomarker identification. Our results show that (i) NBBD is competitive with some of the state-of-the-art feature selection methods including Random Forest Feature Importance (RFFI) scores; and (ii) NBBD is especially effective in reliably identifying IBD biomarkers when the number of data samples available for biomarker discovery is small.