Comparing and Validating Machine Learning Models for Mycobacterium tuberculosis Drug Discovery.
Comparing and Validating Machine Learning Models for Mycobacterium tuberculosis Drug Discovery.
复制标题
DOI:
10.1021/acs.molpharmaceut.8b00083
复制
发表时间:
2018-10-01
影响因子:
4.9
通讯作者:
Ekins S
中科院分区:
文献类型:
--
作者:
Lane T;Russo DP;Zorn KM;Clark AM;Korotcov A;Tkachenko V;Reynolds RC;Perryman AL;Freundlich JS;Ekins S
Tuberculosis is a global health dilemma. In 2016, the WHO reported 10.4 million incidences and 1.7 million deaths. The need to develop new treatments for those infected with Mycobacterium tuberculosis (Mtb) has led to many large-scale phenotypic screens and many thousands of new active compounds identified in vitro. However, with limited funding, efforts to discover new active molecules against Mtb needs to be more efficient. Several computational machine learning approaches have been shown to have good enrichment and hit rates. We have curated small molecule Mtb data and developed new models with a total of 18,886 molecules with activity cut offs of 10 µM, 1 µM and 100 nM. These datasets were used to evaluate different machine learning methods (including deep learning) and metrics and generate predictions for additional molecules published in 2017. One Mtb model, a combined in vitro and in vivo data Bayesian model at a 100 nM activity yielded the following metrics for 5-fold cross validation: Accuracy = 0.88, Precision = 0.22, Recall = 0.91, Specificity = 0.88, Kappa = 0.31, and MCC = 0.41. We have also curated an evaluation set (n = 153 compounds) published in 2017 and when used to test our model it showed the comparable statistics (Accuracy = 0.83, Precision = 0.27, Recall = 1.00, Specificity = 0.81, Kappa = 0.36, and MCC = 0.47). We have also compared these models with additional machine learning algorithms showing Bayesian machine learning models constructed with literature Mtb data generated by different labs generally were equivalent to or outperformed Deep Neural Networks with external test sets. Finally, we have also compared our training and test sets to show they were suitably diverse and different in order to represent useful evaluation sets. Such Mtb machine learning models could help prioritize compounds for testing in vitro and in vivo.
登录
查看更多内容
影响因子:
3.2
作者:
Ekins, Sean;Godbole, Adwait Anand;Nagaraja, Valakunja
通讯作者:
Nagaraja, Valakunja
影响因子:
5.6
作者:
Clark AM;Dole K;Ekins S
通讯作者:
Ekins S
影响因子:
3.7
作者:
Ekins, Sean;Freundlich, Joel S.
通讯作者:
Freundlich, Joel S.
影响因子:
5.6
作者:
Ekins S;Pottorf R;Reynolds RC;Williams AJ;Clark AM;Freundlich JS
通讯作者:
Freundlich JS
影响因子:
3.7
作者:
Ekins, Sean;Freundlich, Joel S.;Hobrath, Judith V.;White, E. Lucile;Reynolds, Robert C.
通讯作者:
Reynolds, Robert C.