Kinome-wide activity modeling from diverse public high-quality data sets.

Kinome-wide activity modeling from diverse public high-quality data sets.
复制标题

DOI:
10.1021/ci300403k
复制
发表时间:
2013-01-28
影响因子:
5.6
通讯作者:
Muskal SM
Muskal SM
中科院分区:
化学2区
文献类型:
--
作者:
Schürer SC;Muskal SM

文献摘要

参考文献

被引文献

相似文献

激酶小分子抑制剂数据的大型语料库可从数千篇期刊文章和专利出版物中获得公共部门研究。这些数据是由许多实验室采用各种各样的测定方法和实验程序产生的。在这里,我们问的问题是如何适用于这些异构数据集预测激酶活性和数据集的哪些特征有助于其效用。我们从激酶知识库(KKB)中访问了近500,000个分子,经过严格的聚合和标准化,生成了180多个不同的数据集,涵盖了人类Kinome的所有主要群体。为了评估数据集的价值,我们生成了数百个分类和回归模型。他们严格的交叉验证和表征表明,如果有最少所需数量的活性化合物或结构活性数据点,大多数激酶靶点具有高度预测性的分类和定量模型。然后,我们将最好的分类器应用于最近在NIH基于集成网络的细胞签名(LINCS)程序库中分析的化合物,并发现分析结果与预测活动的良好一致性。我们的研究结果表明,虽然本质上是异质的,但可访问的数据集非常有价值,非常适合开发高度准确的预测因子,用于实际的全Kinome虚拟筛选应用,并补充实验激酶分析。
Large corpora of kinase small molecule inhibitor data are accessible to public sector research from thousands of journal article and patent publications. These data have been generated employing a wide variety of assay methodologies and experimental procedures by numerous laboratories. Here we ask the question how applicable these heterogeneous datasets are to predict kinase activities and which characteristics of the datasets contribute to their utility. We accessed almost 500,000 molecules from the Kinase Knowledge Base (KKB) and after rigorous aggregation and standardization generated over 180 distinct datasets covering all major groups of the human Kinome. To assess the value of the datasets we generated hundreds of classification and regression models. Their rigorous cross-validation and characterization demonstrated highly predictive classification and quantitative models for the majority of kinase targets if a minimum required number of active compounds or structure-activity data points were available. We then applied the best classifiers to compounds most recently profiled in the NIH Library of Integrated Network-based Cellular Signatures (LINCS) program and found good agreement of profiling results with predicted activities. Our results indicate that, although heterogeneous in nature, the publically accessible datasets are exceedingly valuable and well suited to develop highly accurate predictors for practical Kinome-wide virtual screening applications and to complement experimental kinase profiling.
DOI: 10.1002/cbic.200400109
发表时间: 2005-03-01
期刊: CHEMBIOCHEM
影响因子: 3.2
作者:
Briem, H;Günther, J
通讯作者: Günther, J
DOI: 10.1073/pnas.0708800104
发表时间: 2007-12-18
影响因子: 11.1
作者:
Fedorov, Oleg;Marsden, Brian;Knapp, Stefan
通讯作者: Knapp, Stefan
DOI: 10.1038/nchembio799
发表时间: 2006-07-01
影响因子: 14.8
作者:
Liu, Y;Gray, NS
通讯作者: Gray, NS
DOI: 10.1021/jm8011036
发表时间: 2008-12-25
影响因子: 7.3
作者:
Bamborough, Paul;Drewry, David;Schneider, Klaus
通讯作者: Schneider, Klaus
DOI: 10.1021/jm701021b
发表时间: 2008-03-13
影响因子: 7.3
作者:
Aronov, Alex M.;McClain, Brian;Murcko, Mark A.
通讯作者: Murcko, Mark A.