A CitationRank algorithm inheriting Google technology designed to highlight genes responsible for serious adverse drug reaction

A CitationRank algorithm inheriting Google technology designed to highlight genes responsible for serious adverse drug reaction
复制标题

DOI:
10.1093/bioinformatics/btp369
复制
发表时间:
2009-09-01
期刊:
影响因子:
5.8
通讯作者:
He, Lin
He, Lin
中科院分区:
生物学3区
文献类型:
--
作者:
Yang, Lun;Xu, Langlai;He, Lin

文献摘要

被引文献

相似文献

动机:严重药物不良反应(SADR)是一个紧迫的世界性问题。在没有任何组织良好的面向基因的SADR信息库的情况下,应该构建一个数据库。由于基因对特定 SADR 的重要性不能简单地根据两者在文献中一起引用的频率来定义,因此应设计一种算法根据基因与 SADR 主题的相关性对基因进行排序。 结果:由 Pubmed 提取的基因-SADR 关系组成的 SADR-Gengle 数据库已构建,涵盖六种主要 SADR,即胆汁淤积、耳聋、肌肉毒性、QT 延长、Stevens-Johnson 综合征和扭转型室速德点。 CitationRank算法继承了Google PageRank算法的原理,即当一个基因与其他排名靠前的基因在生物学上相关时,该基因应该排名靠前。该算法在存在外来噪声的情况下恢复 SADR 相关基因方面表现强劲,并且该算法的使用已扩展到对我们数据库中的基因进行排序。用户可以在 Google 类型的系统中浏览基因,其中基因根据其与用户选择的 SADR 主题的相关性降序排列。该数据库还为用户提供了可视化的基因-基因知识链网络,帮助用户在浏览这些网络的同时将其面向基因的知识链系统化。
Motivation: Serious adverse drug reaction (SADR) is an urgent, world-wide problem. In the absence of any well-organized gene-oriented SADR information pool, a database should be constructed. Since the importance of a gene to a particular SADR cannot simply be defined in terms of how frequently the two are cited together in the literature, an algorithm should be devised to sort genes according to their relevance to the SADR topics.Results: The SADR-Gengle database, which is made up of gene-SADR relationships extracted from Pubmed, has been constructed, covering six major SADRs, namely cholestasis, deafness, muscle toxicity, QT prolongation, Stevens-Johnson syndrome and torsades de points. The CitationRank algorithm, which inherits the principle of the Google PageRank algorithm that a gene should be highly ranked when biologically related to other highly ranked genes, is devised. The algorithm performs robustly in recovering SADR-related genes in the presence of extraneous noise, and the use of the algorithm has been extended to sorting genes in our database. Users can browse genes in a Google-type system where genes are ordered according to their descending relevance to the SADR topic selected by the user. The database also provides users with visualized gene-gene knowledge chain networks, helping them to systematize their gene-oriented knowledge chain whilst navigating these networks.