Open Source Bayesian Models. 3. Composite Models for Prediction of Binned Responses.

Open Source Bayesian Models. 3. Composite Models for Prediction of Binned Responses.
复制标题

DOI:
10.1021/acs.jcim.5b00555
复制
发表时间:
2016-02-22
影响因子:
5.6
通讯作者:
Ekins S
Ekins S
中科院分区:
化学2区
文献类型:
--
作者:
Clark AM;Dole K;Ekins S

文献摘要

被引文献

相似文献

贝叶斯模型构建的结构衍生的指纹已成为一种流行的和有用的方法,用于药物发现研究时,应用于生物活性的测量,可以有效地分为活性或非活性。结果可用于根据其活性概率对候选结构进行排名,并且当使用基于结构的指纹时,这种排名受益于高度的可解释性,从而使结果在化学上直观。除了选择活动阈值外,构建贝叶斯模型的速度很快,并且几乎不需要参数或用户干预。该方法也没有遭受急性过度训练的问题,如定量结构活性关系或定量结构性质关系(QSAR/QSPR)。这使得它非常适合于独立于用户专业知识或训练数据的先验知识的自动化工作流程。我们现在描述一种新的方法,用于创建贝叶斯模型的复合组,以扩展该方法以处理多个状态,而不仅仅是二进制。传入的活动被划分到多个箱中,每个箱覆盖相互排斥的活动范围。对于这些箱中的每一个,创建贝叶斯模型以对化合物是否属于该箱进行建模。使用复合模型分析推定的分子涉及对每个箱进行预测并检查每个分配的相对可能性,例如,最高值获胜。已在从ChEMBL v20中提取的数百个数据集以及ADME/Tox和生物活性的经验证数据集上对该方法进行了评价。
Bayesian models constructed from structure-derived fingerprints have been a popular and useful method for drug discovery research when applied to bioactivity measurements that can be effectively classified as active or inactive. The results can be used to rank candidate structures according to their probability of activity, and this ranking benefits from the high degree of interpretability when structure-based fingerprints are used, making the results chemically intuitive. Besides selecting an activity threshold, building a Bayesian model is fast and requires few or no parameters or user intervention. The method also does not suffer from such acute overtraining problems as quantitative structure–activity relationships or quantitative structure–property relationships (QSAR/QSPR). This makes it an approach highly suitable for automated workflows that are independent of user expertise or prior knowledge of the training data. We now describe a new method for creating a composite group of Bayesian models to extend the method to work with multiple states, rather than just binary. Incoming activities are divided into bins, each covering a mutually exclusive range of activities. For each of these bins, a Bayesian model is created to model whether or not the compound belongs in the bin. Analyzing putative molecules using the composite model involves making a prediction for each bin and examining the relative likelihood for each assignment, for example, highest value wins. The method has been evaluated on a collection of hundreds of data sets extracted from ChEMBL v20 and validated data sets for ADME/Tox and bioactivity.