Impact of Gene Biomarker Discovery Tools Based on Protein-Protein Interaction and Machine Learning on Performance of Artificial Intelligence Models in Predicting Clinical Stages of Breast Cancer

Impact of Gene Biomarker Discovery Tools Based on Protein-Protein Interaction and Machine Learning on Performance of Artificial Intelligence Models in Predicting Clinical Stages of Breast Cancer
复制标题

DOI:
10.1007/s12539-020-00390-8
复制
发表时间:
2020-09-10
影响因子:
4.8
通讯作者:
Dastmalchi, Siavoush
Dastmalchi, Siavoush
中科院分区:
生物学3区
文献类型:
--
作者:
Amjad, Elham;Asnaashari, Solmaz;Dastmalchi, Siavoush

文献摘要

被引文献

相似文献

乳腺癌作为威胁妇女生命的常见疾病之一,已引起世界各国临床和生物医学研究者的高度重视。基于基因组的研究沿着其注册的GEO数据集在文献中是常见的。由于已经开发了几种用于分析和鉴定基因生物标志物的方法,因此有必要评估它们的鲁棒性。在这项研究中,三种众所周知的生物标志物鉴定方法(即,使用生物标记物(BioDiscML、MCODE和BioDiscOne)以鉴定潜在的生物标记物。然后,使用基于所识别的生物标志物集开发的非线性分类模型对这些方法进行排名和评估。将由GSE 124647、GSE 124646和GSE 15852组成的组合BC微阵列数据集用作训练集,并且将两个测试数据集GSE 15852和GSE 25066用于训练模型的性能测量。所提出的模型的验证进行了内部(留一法,五倍和十倍交叉验证,随机抽样,测试训练集)和外部(测试集测试)。结果显示,根据曲线下面积(AUC)、准确率、F1评分、精确度和召回率指标,QuantiterOne、MCODE和BioDiscML工具分别排名第一、第二和第三。总体而言,可以得出结论,在验证生物标志物鉴定方法的同时,应同时考虑基因生物标志物在其生物学方面的描述性值(已通过给定方法确定)和基于所鉴定的基因生物标志物开发的模型的预测能力。
Breast cancer, as one of the most common diseases threatening the women's life, has attracted serious attention of the clinical and biomedical researchers worldwide. The genome-based studies along with their registered GEO datasets are frequent in the literature. Since several methodologies have been developed for analyzing and identifying gene biomarkers, it is necessary to evaluate their robustness. In this study, three well-known biomarker identification methods (i.e., ClusterOne, MCODE, and BioDiscML) were employed in order to identify the potential biomarkers. Then, the methods were ranked and evaluated using nonlinear classification models developed based on the identified sets of biomarkers. A combined BC microarray dataset consisting of GSE124647, GSE124646, and GSE15852 was used as training set, and two test datasets, GSE15852 and GSE25066, were used for the performance measurement of the trained models. The validation of the proposed models was carried out internally (leave-one-out, fivefold and tenfold cross-validation, random sampling, test on training set) and externally (test on test set). The results showed that ClusterOne, MCODE, and BioDiscML tools ranked first, second, and third, respectively, based on the area under the curve (AUC), accuracy, F1 score, precision, and recall metrics. Overall, it can be concluded that the descriptive values of gene biomarkers in terms of their biological aspects that have been determined by a given methodology and the predictive power of the models developed based on the identified gene biomarkers should be considered simultaneously while validating the biomarker identification approaches.