STarFish: A Stacked Ensemble Target Fishing Approach and its Application to Natural Products

STarFish: A Stacked Ensemble Target Fishing Approach and its Application to Natural Products
复制标题

DOI:
10.1021/acs.jcim.9b00489
复制
发表时间:
2019-11-01
影响因子:
5.6
通讯作者:
Fuchs, James R.
Fuchs, James R.
中科院分区:
化学2区
文献类型:
--
作者:
Cockroft, Nicholas T.;Cheng, Xiaolin;Fuchs, James R.

文献摘要

被引文献

相似文献

目标钓鱼是识别生物活性小分子的蛋白质目标的过程。要通过实验来实现这一点,需要投入大量的时间和资源,而这可以通过可靠的计算目标捕捞模型来加速。由于大量公共生物活动数据的可用性不断增加,使用机器学习开发计算目标捕鱼模型在过去几年中变得非常流行。不幸的是,此类模型对天然产品的适用性和性能尚未得到全面评估。部分原因是与合成化合物相比,天然产物的生物活性数据相对缺乏。此外,通常用于训练此类模型的数据库不会注释哪些化合物是天然产物,这使得基准测试集的收集变得困难。为了解决这一知识差距,通过交叉引用 20 个公开的天然产物数据库与生物活性数据库 ChEMBL,生成了由天然产物结构及其相关蛋白质靶标组成的数据集。该数据集包含 1943 个独特化合物和 1023 个独特目标的 5589 个化合物-目标对。包含 88,728 个独特化合物和 1907 个独特目标的 107,190 个化合物-目标对的合成数据集用于训练 k 最近邻、随机森林和多层感知器模型。每个模型的预测性能通过分层 10 倍交叉验证和新收集的天然产品数据集的基准测试来评估。在交叉验证过程中,每个模型都表现出很强的性能,接收者操作特征面积 (AUROC) 得分范围为 0.94 至 0.99,玻尔兹曼增强接收者操作特征辨别 (BEDROC) 得分为 0.89 至 0.94。在天然产品数据集上进行测试时,性能急剧下降,AUROC 分数从 0.70 到 0.85,BEDROC 分数从 0.43 到 0.59。然而,使用逻辑回归作为元分类器来组合模型预测的模型堆叠方法的实施,极大地提高了正确预测天然产物蛋白质目标的能力,并将 AUROC 得分提高到 0.94,BEDROC 得分提高到 0.73。该堆叠模型被部署为一个名为 STarFish 的 Web 应用程序,并且已可用于帮助天然产品的目标识别。
Target fishing is the process of identifying the protein target of a bioactive small molecule. To do so experimentally requires a significant investment of time and resources, which can be expedited with a reliable computational target fishing model. The development of computational target fishing models using machine learning has become very popular over the last several years because of the increased availability of large amounts of public bioactivity data. Unfortunately, the applicability and performance of such models for natural products has not yet been comprehensively assessed. This is, in part, due to the relative lack of bioactivity data available for natural products compared to synthetic compounds. Moreover, the databases commonly used to train such models do not annotate which compounds are natural products, which makes the collection of a benchmarking set difficult. To address this knowledge gap, a data set composed of natural product structures and their associated protein targets was generated by cross-referencing 20 publicly available natural product databases with the bioactivity database ChEMBL. This data set contains 5589 compound-target pairs for 1943 unique compounds and 1023 unique targets. A synthetic data set comprising 107 190 compound-target pairs for 88 728 unique compounds and 1907 unique targets was used to train k-nearest neighbors, random forest, and multilayer perceptron models. The predictive performance of each model was assessed by stratified 10-fold cross-validation and benchmarking on the newly collected natural product data set. Strong performance was observed for each model during cross-validation with area under the receiver operating characteristic (AUROC) scores ranging from 0.94 to 0.99 and Boltzmann-enhanced discrimination of receiver operating characteristic (BEDROC) scores from 0.89 to 0.94. When tested on the natural product data set, performance dramatically decreased with AUROC scores ranging from 0.70 to 0.85 and BEDROC scores from 0.43 to 0.59. However, the implementation of a model stacking approach, which uses logistic regression as a meta-classifier to combine model predictions, dramatically improved the ability to correctly predict the protein targets of natural products and increased the AUROC score to 0.94 and BEDROC score to 0.73. This stacked model was deployed as a web application, called STarFish, and has been made available for use to aid in target identification for natural products.