Mapping between databases of compounds and protein targets.

Mapping between databases of compounds and protein targets.
复制标题

DOI:
10.1007/978-1-61779-965-5_8
复制
发表时间:
2012-01-01
期刊:
Methods in molecular biology (Clifton, N.J.)
影响因子:
--
通讯作者:
Southan, Christopher
Southan, Christopher
中科院分区:
其他
文献类型:
--
作者:
Muresan, Sorel;Sitzmann, Markus;Southan, Christopher

文献摘要

被引文献

相似文献

提供生物活性化合物与其蛋白质靶标之间联系的数据库在药物发现和化学生物学中变得越来越重要。它们一方面通过化学结构加入不断扩大的化学信息学领域,另一方面通过序列加入生物信息学领域。然而,如果没有明确的内容比较,就很难评估数据库的相对效用。我们通过比较在化学结构和蛋白质水平上对生物活性化学(ChEMBL、DrugBank、人类代谢组数据库和治疗靶点数据库)有不同关注点的资源,举例说明了一种方法。我们使用 NCI/CADD 结构标识符比较了不同表征严格性下的化合物集。化学含量的重叠和独特性可以在不同的数据捕获策略的背景下得到广泛的解释。然而,我们记录了明显的异常情况,例如代谢物和药物数据库之间存在许多共同化合物。我们还通过 UniProt 蛋白质标识符比较了化合物映射的序列内容。虽然这些通常在各个数据库的背景下也可以解释,但我们发现了覆盖范围和所使用的支持数据类型的差异。例如,DrugBank 和治疗目标数据库之间的目标概念应用不同。在 ChEMBL 中,除了药物靶点本身之外,它还包括化学生物学和物种直向同源物交叉筛选的更广泛的映射。我们的分析不仅应该帮助用户利用这四种高价值资源之间的协同作用,而且还可以评估化学和生物学领域其他数据库的效用。
Databases that provide links between bioactive compounds and their protein targets are increasingly important in drug discovery and chemical biology. They join the expanding universes of cheminformatics via chemical structures on the one hand and bioinformatics via sequences on the other. However, it is difficult to assess the relative utility of databases without the explicit comparison of content. We have exemplified an approach to this by comparing resources that each has a different focus on bioactive chemistry (ChEMBL, DrugBank, Human Metabolome Database, and Therapeutic Target Database) both at the chemical structure and protein levels. We compared the compound sets at different representational stringencies using NCI/CADD Structure Identifiers. The overlap and uniqueness in chemical content can be broadly interpreted in the context of different data capture strategies. However, we recorded apparent anomalies, such as many compounds-in-common between the metabolite and drug databases. We also compared the content of sequences mapped to the compounds via their UniProt protein identifiers. While these were also generally interpretable in the context of individual databases we discerned differences in coverage and the types of supporting data used. For example, the target concept is applied differently between DrugBank and the Therapeutic Target Database. In ChEMBL it encompasses a broader range of mappings from chemical biology and species orthologue cross-screening in addition to drug targets per se. Our analysis should assist users not only in exploiting the synergies between these four high-value resources but also in assessing the utility of other databases at the interface of chemistry and biology.