Hit Dexter 2.0: Machine-Learning Models for the Prediction of Frequent Hitters

Hit Dexter 2.0: Machine-Learning Models for the Prediction of Frequent Hitters
复制标题

DOI:
10.1021/acs.jcim.8b00677
复制
发表时间:
2019-03-01
影响因子:
5.6
通讯作者:
Kirchmair, Johannes
Kirchmair, Johannes
中科院分区:
化学2区
文献类型:
--
作者:
Stork, Conrad;Chen, Ya;Kirchmair, Johannes

文献摘要

被引文献

相似文献

由小分子引起的分析干扰继续对早期药物发现构成重大挑战。已有许多基于规则和基于相似性的方法,允许标记潜在的“行为恶劣的化合物”、“不良行为者”或“滋扰化合物”。这些化合物通常是聚集剂、活性化合物和/或泛分析干扰化合物(Pain),其中许多是经常使用的药物。Hit Dexter是最近引入的一种机器学习方法,它预测频繁的Hit,而不依赖于潜在的物理化学机制(也包括基于“特权支架”的化合物与多个结合位点的结合)。在这里,我们报告了第二代机器学习模型的发展,现在包括初步筛选分析和验证性剂量反应分析。蛋白质序列聚类是为了最大限度减少结构和功能相关蛋白质的过度表达而引入的。该模型正确地将大型独立测试集的化合物分类为(高度)混杂或非混杂,马修斯相关系数(MCC)值高达0.64,接收器工作特征曲线下面积(AUC)值高达0.96。这些模型还被用来表征具有特定生物学和物理化学性质的化合物集合,如暗化学物质、聚合体、来自高通量筛选库的化合物、类药物化合物、批准的药物、潜在的疼痛和天然产品。最有趣的结果之一是,新的HIT Dexter模型预测,在批准的药物中存在大量(高度)混杂化合物。重要的是,个人点击Dexter模型的预测总体上是一致的,并与bAdapple的预测一致,bAdapple是一个为预测频繁点击的人而建立的统计模型。新的Hit Dexter2.0 Web服务可在http://hitdexter2.zbh.uni-hamburg.de,上获得,它不仅提供了对这项工作中提出的所有机器学习模型的用户友好访问,还提供了用于预测聚集体和暗化学物质的基于相似性的方法,以及用于标记频繁点击者和包括不需要的子结构的化合物的可用规则集的全面集合。
Assay interference caused by small molecules continues to pose a significant challenge for early drug discovery. A number of rule-based and similarity-based approaches have been derived that allow the flagging of potentially "badly behaving compounds", "bad actors", or "nuisance compounds". These compounds are typically aggregators, reactive compounds, and/or pan-assay interference compounds (PAINS), and many of them are frequent hitters. Hit Dexter is a recently introduced machine learning approach that predicts frequent hitters independent of the underlying physicochemical mechanisms (including also the binding of compounds based on "privileged scaffolds" to multiple binding sites). Here we report on the development of a second generation of machine learning models which now covers both primary screening assays and confirmatory dose response assays. Protein sequence clustering was newly introduced to minimize the overrepresentation of structurally and functionally related proteins. The models correctly classified compounds of large independent test sets as (highly) promiscuous or nonpromiscuous with Matthews correlation coefficient (MCC) values of up to 0.64 and area under the receiver operating characteristic curve (AUC) values of up to 0.96. The models were also utilized to characterize sets of compounds with specific biological and physicochemical properties, such as dark chemical matter, aggregators, compounds from a high-throughput screening library, drug-like compounds, approved drugs, potential PAINS, and natural products. Among the most interesting outcomes is that the new Hit Dexter models predict the presence of large fractions of (highly) promiscuous compounds among approved drugs. Importantly, predictions of the individual Hit Dexter models are generally in good agreement and consistent with those of Badapple, an established statistical model for the prediction of frequent hitters. The new Hit Dexter 2.0 web service, available at http://hitdexter2.zbh.uni-hamburg.de, not only provides user-friendly access to all machine learning models presented in this work but also to similarity-based methods for the prediction of aggregators and dark chemical matter as well as a comprehensive collection of available rule sets for flagging frequent hitters and compounds including undesired substructures.