Identifying gene-disease associations using centrality on a literature mined gene-interaction network.

Identifying gene-disease associations using centrality on a literature mined gene-interaction network.
复制标题

使用中心性在文献挖掘的基因互动网络上识别基因 - 疾病的关联。

DOI:
10.1093/bioinformatics/btn182
复制
发表时间:
2008-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Radev DR
Radev DR
中科院分区:
其他
文献类型:
--
作者:
Ozgür A;Vu T;Erkan G;Radev DR

文献摘要

参考文献

被引文献

相似文献

动机:了解遗传学在疾病中的作用是生物科学最重要的目标之一。人类基因组计划的完成使这一领域的出版物数量迅速增加。然而,提供从文献中手动提取的信息的策展数据库的覆盖范围有限。另一个挑战是确定疾病相关基因需要艰苦的实验。因此,在实验分析之前预测好的候选基因将节省时间和精力。我们介绍了一种自动的方法,基于文本挖掘和网络分析来预测基因-疾病的关联。我们收集了一组已知的疾病相关基因的初始集,并建立了一个相互作用的网络,自动文献挖掘依赖分析和支持向量机的基础上。我们的假设是,这个疾病特异性网络中的中心基因可能与疾病有关。我们使用度、特征向量、介数和接近中心性度量对网络中的基因进行排序。结果:所提出的方法可以用来提取已知和推断未知的基因-疾病关联。我们评估了前列腺癌的方法。特征向量和度中心性算法具有较高的精度。通过这些方法排名的前20个基因中,共有95%被证实与前列腺癌有关。另一方面,介数和接近中心性预测更多的基因,其与疾病的关系是目前未知的,是实验研究的候选人。可用性:浏览疾病特异性基因相互作用网络的基于网络的系统可在:http://gin.ncibi.org联系:radev@umich.edu
Motivation: Understanding the role of genetics in diseases is one of the most important aims of the biological sciences. The completion of the Human Genome Project has led to a rapid increase in the number of publications in this area. However, the coverage of curated databases that provide information manually extracted from the literature is limited. Another challenge is that determining disease-related genes requires laborious experiments. Therefore, predicting good candidate genes before experimental analysis will save time and effort. We introduce an automatic approach based on text mining and network analysis to predict gene-disease associations. We collected an initial set of known disease-related genes and built an interaction network by automatic literature mining based on dependency parsing and support vector machines. Our hypothesis is that the central genes in this disease-specific network are likely to be related to the disease. We used the degree, eigenvector, betweenness and closeness centrality metrics to rank the genes in the network. Results: The proposed approach can be used to extract known and to infer unknown gene-disease associations. We evaluated the approach for prostate cancer. Eigenvector and degree centrality achieved high accuracy. A total of 95% of the top 20 genes ranked by these methods are confirmed to be related to prostate cancer. On the other hand, betweenness and closeness centrality predicted more genes whose relation to the disease is currently unknown and are candidates for experimental study. Availability: A web-based system for browsing the disease-specific gene-interaction networks is available at: http://gin.ncibi.org Contact: radev@umich.edu
DOI: 10.1186/1471-2105-5-147
发表时间: 2004-10-08
期刊: BMC bioinformatics
影响因子: 3
作者:
Chen H;Sharp BM
通讯作者: Sharp BM
DOI: 10.1093/nar/gkg008
发表时间: 2003-01-01
影响因子: 14.9
作者:
Li, LC;Zhao, H;Dahiya, R
通讯作者: Dahiya, R
DOI: 10.1038/35075138
发表时间: 2001-05-03
期刊: NATURE
影响因子: 64.8
作者:
Jeong, H;Mason, SP;Oltvai, ZN
通讯作者: Oltvai, ZN
DOI: 10.1038/sj.bjc.6600747
发表时间: 2003-01-27
影响因子: 8.8
作者:
通讯作者: --
酵母蛋白相互作用网络中的高比分蛋白。
DOI: 10.1155/jbb.2005.96
发表时间: 2005-06-30
影响因子: --
作者:
Joy, MP;Brock, A;Ingber, DE;Huang, S
通讯作者: Huang, S