Comparing and Validating Machine Learning Models for Mycobacterium tuberculosis Drug Discovery.

Comparing and Validating Machine Learning Models for Mycobacterium tuberculosis Drug Discovery.
复制标题

DOI:
10.1021/acs.molpharmaceut.8b00083
复制
发表时间:
2018-10-01
影响因子:
4.9
通讯作者:
Ekins S
Ekins S
中科院分区:
医学2区
文献类型:
--
作者:
Lane T;Russo DP;Zorn KM;Clark AM;Korotcov A;Tkachenko V;Reynolds RC;Perryman AL;Freundlich JS;Ekins S

文献摘要

参考文献

被引文献

相似文献

结核病是一个全球性的健康难题。2016年,世界卫生组织报告了1040万例发病率和170万例死亡。为结核分枝杆菌(Mtb)感染者开发新治疗方法的需求导致了许多大规模的表型筛选和数千种新的体外活性化合物的鉴定。然而,由于资金有限,发现抗结核病新活性分子的努力需要更加有效。几种计算机器学习方法已被证明具有良好的富集和命中率。我们已经策划了小分子Mtb数据,并开发了总共18,886个分子的新模型,活性截止值为10 µ M,1 µ M和100 nM。这些数据集用于评估不同的机器学习方法(包括深度学习)和指标,并为2017年发表的其他分子生成预测。一个Mtb模型,一种在100 nM活性下的组合的体外和体内数据贝叶斯模型,产生了以下5倍交叉验证的指标:准确度= 0.88,精密度= 0.22,召回率= 0.91,特异性= 0.88,Kappa = 0.31,MCC = 0.41。我们还策划了2017年发表的评估集(n = 153种化合物),当用于测试我们的模型时,它显示了可比的统计数据(准确度= 0.83,精度= 0.27,召回率= 1.00,特异性= 0.81,Kappa = 0.36,MCC = 0.47)。我们还将这些模型与其他机器学习算法进行了比较,结果显示,使用不同实验室生成的文献Mtb数据构建的贝叶斯机器学习模型通常等同于或优于使用外部测试集的深度神经网络。最后,我们还比较了我们的训练集和测试集,以表明它们是适当的多样性和不同的,以代表有用的评估集。这种Mtb机器学习模型可以帮助优先考虑体外和体内测试的化合物。
Tuberculosis is a global health dilemma. In 2016, the WHO reported 10.4 million incidences and 1.7 million deaths. The need to develop new treatments for those infected with Mycobacterium tuberculosis (Mtb) has led to many large-scale phenotypic screens and many thousands of new active compounds identified in vitro. However, with limited funding, efforts to discover new active molecules against Mtb needs to be more efficient. Several computational machine learning approaches have been shown to have good enrichment and hit rates. We have curated small molecule Mtb data and developed new models with a total of 18,886 molecules with activity cut offs of 10 µM, 1 µM and 100 nM. These datasets were used to evaluate different machine learning methods (including deep learning) and metrics and generate predictions for additional molecules published in 2017. One Mtb model, a combined in vitro and in vivo data Bayesian model at a 100 nM activity yielded the following metrics for 5-fold cross validation: Accuracy = 0.88, Precision = 0.22, Recall = 0.91, Specificity = 0.88, Kappa = 0.31, and MCC = 0.41. We have also curated an evaluation set (n = 153 compounds) published in 2017 and when used to test our model it showed the comparable statistics (Accuracy = 0.83, Precision = 0.27, Recall = 1.00, Specificity = 0.81, Kappa = 0.36, and MCC = 0.47). We have also compared these models with additional machine learning algorithms showing Bayesian machine learning models constructed with literature Mtb data generated by different labs generally were equivalent to or outperformed Deep Neural Networks with external test sets. Finally, we have also compared our training and test sets to show they were suitably diverse and different in order to represent useful evaluation sets. Such Mtb machine learning models could help prioritize compounds for testing in vitro and in vivo.
DOI: 10.1016/j.tube.2017.01.005
发表时间: 2017-03-01
期刊: TUBERCULOSIS
影响因子: 3.2
作者:
Ekins, Sean;Godbole, Adwait Anand;Nagaraja, Valakunja
通讯作者: Nagaraja, Valakunja
DOI: 10.1021/acs.jcim.5b00555
发表时间: 2016-02-22
影响因子: 5.6
作者:
Clark AM;Dole K;Ekins S
通讯作者: Ekins S
DOI: 10.1007/s11095-011-0413-x
发表时间: 2011-08-01
影响因子: 3.7
作者:
Ekins, Sean;Freundlich, Joel S.
通讯作者: Freundlich, Joel S.
DOI: 10.1021/ci500077v
发表时间: 2014-04-28
影响因子: 5.6
作者:
Ekins S;Pottorf R;Reynolds RC;Williams AJ;Clark AM;Freundlich JS
通讯作者: Freundlich JS
结合命中率优化的计算方法在结核分枝杆菌中的铅优化。
DOI: 10.1007/s11095-013-1172-7
发表时间: 2014-02
影响因子: 3.7
作者:
Ekins, Sean;Freundlich, Joel S.;Hobrath, Judith V.;White, E. Lucile;Reynolds, Robert C.
通讯作者: Reynolds, Robert C.