Assessing drug target association using semantic linked data.

Assessing drug target association using semantic linked data.
复制标题

DOI:
10.1371/journal.pcbi.1002574
复制
发表时间:
2012
影响因子:
4.3
通讯作者:
Wild DJ
Wild DJ
中科院分区:
生物学2区
文献类型:
--
作者:
Chen B;Ding Y;Wild DJ

文献摘要

参考文献

被引文献

相似文献

化学和生物学领域公共数据量的迅速增加为药物发现的大规模数据挖掘提供了新的机遇。这些异构集的系统集成以及提供对集成集进行数据挖掘的算法将允许研究药物的复杂作用机制。在这项工作中,我们整合并注释了与药物、化合物、蛋白质靶点、疾病、副作用和途径相关的公共数据集中的数据,构建了一个由超过 290,000 个节点和 720,000 个边组成的语义链接网络。我们开发了一个统计模型,根据药物靶点对与其他链接对象的关系来评估药物靶点对的关联性。验证实验表明该模型可以高精度地正确识别已知的直接药物靶点对。间接药物靶点对(例如改变基因表达水平的药物)也被识别,但不如直接药物靶点对那么强烈。我们进一步计算了来自 10 个疾病领域的 157 种药物与 1683 个人类目标的关联分数,并使用分数矩阵测量了它们的相似性。相似性网络表明,来自同一疾病领域的药物往往以结构相似性无法捕获的方式聚集在一起,并确定了几种潜在的新药物配对。因此,这项工作为现有药物靶点预测算法提供了一种新颖的、经过验证的替代方案。该网络服务可免费获取:http://chem2bio2rdf.org/slap。现代药物发现需要了解化学基因组学、化合物和药物与体内多种蛋白质靶标和基因的复杂相互作用。与此类关系相关的大量数据存在于可公开访问的数据集中,但它们是孤立的,因此不可能以集成的方式使用。在这项工作中,我们整合并语义注释了来自广泛数据库的大量公共数据,包括化合物-基因、药物-药物、蛋白质-蛋白质、药物-副作用等,以创建与化合物和蛋白质靶标相关的复杂相互作用网络。我们开发了一种称为语义链接关联预测(SLAP)的统计算法,用于预测该数据网络中的“缺失链接”:即没有实验数据但在考虑到该组中存在的其他关系的情况下在统计上可能的复合目标交互。我们提出的验证实验表明该方法具有很高的准确性,并且还演示了如何使用它来创建药物相似性网络以预测现有药物的新适应症。
The rapidly increasing amount of public data in chemistry and biology provides new opportunities for large-scale data mining for drug discovery. Systematic integration of these heterogeneous sets and provision of algorithms to data mine the integrated sets would permit investigation of complex mechanisms of action of drugs. In this work we integrated and annotated data from public datasets relating to drugs, chemical compounds, protein targets, diseases, side effects and pathways, building a semantic linked network consisting of over 290,000 nodes and 720,000 edges. We developed a statistical model to assess the association of drug target pairs based on their relation with other linked objects. Validation experiments demonstrate the model can correctly identify known direct drug target pairs with high precision. Indirect drug target pairs (for example drugs which change gene expression level) are also identified but not as strongly as direct pairs. We further calculated the association scores for 157 drugs from 10 disease areas against 1683 human targets, and measured their similarity using a score matrix. The similarity network indicates that drugs from the same disease area tend to cluster together in ways that are not captured by structural similarity, with several potential new drug pairings being identified. This work thus provides a novel, validated alternative to existing drug target prediction algorithms. The web service is freely available at: http://chem2bio2rdf.org/slap. Modern drug discovery requires the understanding of chemogenomics, the complex interaction of chemical compounds and drugs with a wide variety of protein target and genes in the body. A large amount of data pertaining to such relationships exists in publicly-accessible datasets but it is siloed and thus impossible to use in an integrated fashion. In this work we have integrated and semantically annotated a large amount of public data from a wide range of databases, including compound-gene, drug-drug, protein-protein, drug-side effects and so on, to create a complex network of interactions relating to compounds and protein targets. We developed a statistical algorithm called Semantic Link Association Prediction (SLAP) for predicting “missing links” in this data network: i.e. compound-target interactions for which there is no experimental data but which are statistically probable given the other relationships that exist in this set. We present validation experiments which show this method works with a high degree of accuracy, and also demonstrate how it can be used to create a drug similarity network to make predictions of new indications for existing drugs.
DOI: 10.1186/1471-2105-11-255
发表时间: 2010-05-17
期刊: BMC bioinformatics
影响因子: 3
作者:
Chen B;Dong X;Jiao D;Wang H;Zhu Q;Ding Y;Wild DJ
通讯作者: Wild DJ
DOI: 10.1093/nar/gkp937
发表时间: 2010-01
影响因子: 14.9
作者:
Kuhn M;Szklarczyk D;Franceschini A;Campillos M;von Mering C;Jensen LJ;Beyer A;Bork P
通讯作者: Bork P
DOI: 10.1371/journal.pcbi.1000937
发表时间: 2010-09-01
影响因子: 4.3
作者:
Ferreira, Joao D.;Couto, Francisco M.
通讯作者: Couto, Francisco M.
DOI: 10.1126/science.1158140
发表时间: 2008-07-11
期刊: SCIENCE
影响因子: 56.9
作者:
Campillos, Monica;Kuhn, Michael;Bork, Peer
通讯作者: Bork, Peer
DOI: 10.1371/journal.pcbi.1000423
发表时间: 2009-07
影响因子: 4.3
作者:
Kinnings SL;Liu N;Buchmeier N;Tonge PJ;Xie L;Bourne PE
通讯作者: Bourne PE