Multiple Machine Learning Comparisons of HIV Cell-based and Reverse Transcriptase Data Sets.

Multiple Machine Learning Comparisons of HIV Cell-based and Reverse Transcriptase Data Sets.
复制标题

DOI:
10.1021/acs.molpharmaceut.8b01297
复制
发表时间:
2019-04-01
影响因子:
4.9
通讯作者:
Ekins S
Ekins S
中科院分区:
医学2区
文献类型:
--
作者:
Zorn KM;Lane TR;Russo DP;Clark AM;Makarov V;Ekins S

文献摘要

参考文献

被引文献

相似文献

人体免疫缺陷病毒(艾滋病毒)每年造成100多万人死亡,并在许多国家产生巨大的经济影响。第一类被批准的药物是核苷类逆转录酶抑制剂。新一代逆转录酶抑制剂已变得对HIV耐药株敏感,因此迫切需要替代品。我们最近率先使用贝叶斯机器学习来生成具有公共数据的模型,以识别针对不同疾病靶标进行测试的新化合物。目前的研究使用了NIAID ChemDB HIV,结核病感染和结核病治疗数据库进行机器学习研究。我们从HIV-1野生型细胞和逆转录酶(RT)DNA聚合酶抑制试验中收集和清理数据。该数据库中具有≤ 1μM HIV-1 RT DNA聚合酶活性抑制和基于细胞的HIV-1抑制的化合物具有相关性(Pearson r = 0.44,n = 1137,p < 0.0001)。使用多种机器学习方法(Bernoulli Naive Bayes,AdaBoost决策树,随机森林,支持向量分类,k最近邻和深度神经网络以及共识方法)训练模型,然后比较它们的预测能力。我们对不同机器学习方法的比较表明,支持向量分类,深度学习和共识通常是可比的,并且使用五重交叉验证和使用24个训练和测试集组合彼此没有显著差异。这项研究表明,与我们之前针对各种目标的研究结果一致,使用多个数据集进行训练和测试并没有证明支持向量机和深度神经网络之间存在显着差异。
The human immunodeficiency virus (HIV) causes over a million deaths every year and has a huge economic impact in many countries. The first class of drugs approved were nucleoside reverse transcriptase inhibitors. A newer generation of reverse transcriptase inhibitors have become susceptible to drug resistant strains of HIV, and hence alternatives are urgently needed. We have recently pioneered the use of Bayesian machine learning to generate models with public data to identify new compounds for testing against different disease targets. The current study has used the NIAID ChemDB HIV, Opportunistic Infection and Tuberculosis Therapeutics Database for machine learning studies. We curated and cleaned data from HIV-1 wild-type cell-based and reverse transcriptase (RT) DNA polymerase inhibition assays. Compounds from this database with ≤ 1μM HIV-1 RT DNA polymerase activity inhibition and cell-based HIV-1 inhibition are correlated (Pearson r = 0.44, n = 1137, p < 0.0001). Models were trained using multiple machine learning approaches (Bernoulli Naive Bayes, AdaBoost Decision Tree, Random Forest, support vector classification, k-Nearest Neighbors, and deep neural networks as well as consensus approaches) and then their predictive abilities were compared. Our comparison of different machine learning methods demonstrated that support vector classification, deep learning and a consensus were generally comparable and not significantly different from each other using five-fold cross validation and using 24 training and test set combinations. This study demonstrates findings in line with our previous studies for various targets that training and testing with multiple datasets does not demonstrate a significant difference between support vector machine and deep neural networks.
DOI: 10.1021/jm501908a
发表时间: 2015-03-26
影响因子: 7.3
作者:
Frey KM;Puleo DE;Spasov KA;Bollini M;Jorgensen WL;Anderson KS
通讯作者: Anderson KS
DOI: 10.1039/c3mb70218a
发表时间: 2014-01-01
影响因子: --
作者:
Jain (Pancholi), Nilanjana;Gupta, Swagata;Sapre, Nitin S.
通讯作者: Sapre, Nitin S.
DOI: 10.1016/j.bmcl.2013.12.070
发表时间: 2014-02-01
影响因子: 2.7
作者:
Cote, Bernard;Burch, Jason D.;Ducharme, Yves
通讯作者: Ducharme, Yves
DOI: 10.1371/journal.pntd.0003878
发表时间: 2015
影响因子: 3.8
作者:
Ekins S;de Siqueira-Neto JL;McCall LI;Sarker M;Yadav M;Ponder EL;Kallel EA;Kellar D;Chen S;Arkin M;Bunin BA;McKerrow JH;Talcott C
通讯作者: Talcott C
DOI: 10.1016/b978-0-12-405880-4.00009-3
发表时间: 2013-01-01
期刊: ANTIVIRAL AGENTS
影响因子: --
作者:
De Clercq, Erik
通讯作者: De Clercq, Erik