Understanding false positives in reporter gene assays: in silico chemogenomics approaches to prioritize cell-based HTS data

Understanding false positives in reporter gene assays: in silico chemogenomics approaches to prioritize cell-based HTS data
复制标题

DOI:
10.1021/ci6005504
复制
发表时间:
2007-07-01
影响因子:
5.6
通讯作者:
Glick, Meir
Glick, Meir
中科院分区:
化学2区
文献类型:
--
作者:
Crisman, Thomas J.;Parker, Christian N.;Glick, Meir

文献摘要

被引文献

相似文献

高通量筛选(HTS)数据通常是有噪声的,包含假阳性和假阴性。因此,仔细分类和优先级的主要打击名单可以节省时间和金钱,通过识别潜在的假阳性之前,招致后续费用。特别令人关注的是基于细胞的报告基因测定(RGA),其中命中数可能高得令人望而却步,以至于无法手动仔细检查以清除错误数据。基于从RGA中测试的65万种化合物的化学结构构建的统计模型,我们创建了“频繁命中”模型,可以优先考虑潜在的假阳性。此外,我们跟踪了基于计算机模拟目标预测的化学结构的频繁命中评估,以假设观察到的“脱靶”反应的机制。据观察,已知频繁击球者的预测细胞靶标与诸如细胞毒性的不良效应相关。更具体地说,最常见的预测目标涉及细胞凋亡和细胞分化,包括激酶,拓扑异构酶和蛋白磷酸酶。使用模型预测的160种额外的药物样化合物对基于机制的频繁命中假设进行了检验,这些药物样化合物在RGA中具有非特异性活性。这种验证是成功的(显示50%的命中率,而正常命中率低至2%),它表明了计算模型在理解化学结构和生物功能之间复杂关系方面的能力。
High throughput screening (HTS) data is often noisy, containing both false positives and negatives. Thus, careful triaging and prioritization of the primary hit list can save time and money by identifying potential false positives before incurring the expense of followup. Of particular concern are cell-based reporter gene assays (RGAs) where the number of hits may be prohibitively high to be scrutinized manually for weeding out erroneous data. Based on statistical models built from chemical structures of 650 000 compounds tested in RGAs, we created "frequent hitter" models that make it possible to prioritize potential false positives. Furthermore, we followed up the frequent hitter evaluation with chemical structure based in silico target predictions to hypothesize a mechanism for the observed "off target" response. It was observed that the predicted cellular targets for the frequent hitters were known to be associated with undesirable effects such as cytotoxicity. More specifically, the most frequently predicted targets relate to apoptosis and cell differentiation, including kinases, topoisomerases, and protein phosphatases. The mechanism-based frequent hitter hypothesis was tested using 160 additional druglike compounds predicted by the model to be nonspecific actives in RGAs. This validation was successful (showing a 50% hit rate compared to a normal hit rate as low as 2%), and it demonstrates the power of computational models toward understanding complex relations between chemical structure and biological function.