Analysis of in vitro bioactivity data extracted from drug discovery literature and patents: Ranking 1654 human protein targets by assayed compounds and molecular scaffolds.

Analysis of in vitro bioactivity data extracted from drug discovery literature and patents: Ranking 1654 human protein targets by assayed compounds and molecular scaffolds.
复制标题

DOI:
10.1186/1758-2946-3-14
复制
发表时间:
2011-05-13
影响因子:
8.6
通讯作者:
Muresan S
Muresan S
中科院分区:
化学2区
文献类型:
--
作者:
Southan C;Boppana K;Jagarlapudi SA;Muresan S

文献摘要

参考文献

被引文献

相似文献

自从2002年霍普金斯和格鲁姆发表了经典的可药物基因组综述以来,已经有许多出版物更新了假设的和成功的人类药物目标统计数据。然而,定义这两个极端之间区域的研究目标清单很少,因为在必要的规模上整理已发表的信息是一项挑战。我们通过查询数据库来解决这个问题,这些数据库由专家管理,从过去30年的专利和期刊论文中提取生物活性数据。从超过27,000份文件的子集中,我们提取了一组化合物与靶标的关系,用于1736种人类蛋白质和1654种基因标识符的生化体外结合型分析数据。它们与来自823,179种独特化学结构的1,671,951种化合物记录相关联。该分布表明,每个目标化合物的平均值为964个,最大值为42,869个(因子Xa)。该列表包括非目标、失败目标和交叉筛选目标。前278个最活跃的目标覆盖了90%的化合物。我们通过确定分子框架和支架的数量进一步研究了靶标排序。这些与化合物计数进行比较,作为每个目标基础上化学多样性的替代措施。本工作生成的每蛋白化合物列表(作为补充文件提供)代表了已发表数据定义的人类药物靶点景观的主要比例。我们通过分析化合物的数量来补充简单的排名,并通过分子拓扑来进行额外的排名。这些显示了显著的差异,并提供了化学可追溯性的补充评估。
Since the classic Hopkins and Groom druggable genome review in 2002, there have been a number of publications updating both the hypothetical and successful human drug target statistics. However, listings of research targets that define the area between these two extremes are sparse because of the challenges of collating published information at the necessary scale. We have addressed this by interrogating databases, populated by expert curation, of bioactivity data extracted from patents and journal papers over the last 30 years. From a subset of just over 27,000 documents we have extracted a set of compound-to-target relationships for biochemical in vitro binding-type assay data for 1,736 human proteins and 1,654 gene identifiers. These are linked to 1,671,951 compound records derived from 823,179 unique chemical structures. The distribution showed a compounds-per-target average of 964 with a maximum of 42,869 (Factor Xa). The list includes non-targets, failed targets and cross-screening targets. The top-278 most actively pursued targets cover 90% of the compounds. We further investigated target ranking by determining the number of molecular frameworks and scaffolds. These were compared to the compound counts as alternative measures of chemical diversity on a per-target basis. The compounds-per-protein listing generated in this work (provided as a supplementary file) represents the major proportion of the human drug target landscape defined by published data. We supplemented the simple ranking by the number of compounds assayed with additional rankings by molecular topology. These showed significant differences and provide complementary assessments of chemical tractability.
DOI: 10.1038/nrd2445
发表时间: 2007-11-01
影响因子: 120.1
作者:
Leeson, Paul D.;Springthorpe, Brian
通讯作者: Springthorpe, Brian
DOI: 10.1007/978-1-60761-274-2_6
发表时间: 2009-01-01
期刊: CHEMOGENOMICS: METHODS AND APPLICATIONS
影响因子: --
作者:
Jagarlapudi, Sarma A. R. P.;Kishan, K. V. Radha
通讯作者: Kishan, K. V. Radha
DOI: 10.1002/minf.200900075
发表时间: 2010-01-12
影响因子: 3.6
作者:
Broccatelli, Fabio;Carosati, Emanuele;Cruciani, Gabriele;Oprea, Tudor I.
通讯作者: Oprea, Tudor I.
DOI: 10.1093/nar/gkq1237
发表时间: 2011-01
影响因子: 14.9
作者:
Maglott D;Ostell J;Pruitt KD;Tatusova T
通讯作者: Tatusova T
DOI: 10.1002/cmdc.200900419
发表时间: 2010-02-01
期刊: CHEMMEDCHEM
影响因子: 3.4
作者:
Hu, Ye;Bajorath, Juergen
通讯作者: Bajorath, Juergen