Comparing Multiple Machine Learning Algorithms and Metrics for Estrogen Receptor Binding Prediction.

Comparing Multiple Machine Learning Algorithms and Metrics for Estrogen Receptor Binding Prediction.
复制标题

DOI:
10.1021/acs.molpharmaceut.8b00546
复制
发表时间:
2018-10-01
影响因子:
4.9
通讯作者:
Ekins S
Ekins S
中科院分区:
医学2区
文献类型:
--
作者:
Russo DP;Zorn KM;Clark AM;Zhu H;Ekins S

文献摘要

参考文献

被引文献

相似文献

许多破坏内分泌功能的化学物质与各种不利的生物学结果有关。然而,使用体外或体内方法筛选内分泌干扰既昂贵又耗时。由于更大的训练集、更强的计算能力和先进的机器学习算法(如多层人工神经网络),定量结构-活动关系模型等计算方法变得更加可靠。机器学习模型可用于预测化合物的内分泌干扰能力,如与雌激素受体(ER)结合,并允许优先排序和进一步测试。在这项工作中,对使用内部化学信息学软件(Assay Central)管理的公共数据集进行了多种机器学习算法、化学空间和ER结合评估指标的详尽比较。建模中使用的化学特征包括二进制指纹(ECFP6、FCFP6、ToxPrint或MACCS密钥)和来自RDKit的连续分子描述符。每个特征集都经过经典的机器学习算法(伯努利朴素贝叶斯,AdaBoost决策树,随机森林,支持向量机)和深度神经网络(DNN)。使用各种指标对模型进行评估:召回率、精度、f1评分、准确度、受试者工作特征曲线下面积、Cohen’s Kappa和Matthews相关系数。对于训练集内化合物的预测,DNN具有较高的准确率;然而,在五倍交叉验证和外部测试集预测中,DNN和大多数经典机器学习模型的表现相似,无论使用的数据集或分子描述符如何。我们还使用秩归一化分数作为每个机器学习方法的性能标准,当按度量或按数据集排序时,Random Forest在评估集上表现最好。这些结果表明,经典的机器学习算法可能足以开发高质量的ER活动预测模型。
Many chemicals that disrupt endocrine function have been linked to a variety of adverse biological outcomes. However, screening for endocrine disruption using in vitro or in vivo approaches is costly and time-consuming. Computational methods, e.g. Quantitative Structure-Activity Relationship models, have become more reliable due to bigger training sets, increased computing power, and advanced machine learning algorithms such as multi-layered Artificial Neural Networks. Machine learning models can be used to predict compounds for endocrine disrupting capabilities such as binding to the estrogen receptor (ER) and allow for prioritization and further testing. In this work an exhaustive comparison of multiple machine learning algorithms, chemical spaces, and evaluation metrics for ER binding was performed on public datasets curated using in-house cheminformatics software (Assay Central). Chemical features utilized in modeling consisted of binary fingerprints (ECFP6, FCFP6, ToxPrint or MACCS keys) and continuous molecular descriptors from RDKit. Each feature set was subjected to classic machine learning algorithms (Bernoulli Naive Bayes, AdaBoost Decision Tree, Random Forest, Support Vector Machine) and deep neural networks (DNN). Models were evaluated using a variety of metrics: Recall, Precision, F1-Score, Accuracy, Area Under the Receiver Operating Characteristic Curve, Cohen’s Kappa, and Matthews Correlation Coefficient. For predicting compounds within the training set, DNN has higher Accuracy than other methods; however, in five-fold cross validation and external test set predictions, DNN and most classic machine learning models perform similarly regardless of dataset or molecular descriptors used. We have also used the rank normalized scores as a performance-criteria for each machine learning method and Random Forest performed best on the evaluation set when ranked by metric or by datasets. These results suggest classic machine learning algorithms may be sufficient to develop high quality predictive models of ER activity.
DOI: 10.1038/srep05664
发表时间: 2014-07-11
期刊: Scientific reports
影响因子: 4.6
作者:
Huang R;Sakamuru S;Martin MT;Reif DM;Judson RS;Houck KA;Casey W;Hsieh JH;Shockley KR;Ceger P;Fostel J;Witt KL;Tong W;Rotroff DM;Zhao T;Shinn P;Simeonov A;Dix DJ;Austin CP;Kavlock RJ;Tice RR;Xia M
通讯作者: Xia M
DOI: 10.1080/1062936032000169642
发表时间: 2004-02-01
影响因子: 3
作者:
Asikainen, AH;Ruuskanen, J;Tuppurainen, KA
通讯作者: Tuppurainen, KA
DOI: 10.1021/acs.jcim.5b00555
发表时间: 2016-02-22
影响因子: 5.6
作者:
Clark AM;Dole K;Ekins S
通讯作者: Ekins S
DOI: 10.1038/331091a0
发表时间: 1988-01-07
期刊: NATURE
影响因子: 64.8
作者:
GIGUERE, V;YANG, N;EVANS, RM
通讯作者: EVANS, RM
DOI: 10.1208/s12248-018-0210-0
发表时间: 2018-03-30
期刊: The AAPS journal
影响因子: --
作者:
Jing Y;Bian Y;Hu Z;Wang L;Xie XQ
通讯作者: Xie XQ