GS2: an efficiently computable measure of GO-based similarity of gene sets.

GS2: an efficiently computable measure of GO-based similarity of gene sets.
复制标题

DOI:
10.1093/bioinformatics/btp128
复制
发表时间:
2009-05-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Nakhleh L
Nakhleh L
中科院分区:
其他
文献类型:
--
作者:
Ruths T;Ruths D;Nakhleh L

文献摘要

参考文献

被引文献

相似文献

动机:基因组规模数据集的日益可用性吸引了越来越多的注意力,用于自动推断基因及其产物之间的功能相似性的计算方法的发展。一类这样的方法基于基因本体(GO)中基因的距离来测量基因的功能相似性。为了测量基因集合的功能相关性,这些测量考虑集合中的每对基因,并且计算所有成对距离的平均值。然而,随着更多的数据变得可用并且用于分析的基因组变得更大,这种基于配对的计算变得令人望而却步。结果如下:在这篇文章中,我们提出了GS 2(基于GO的基因集相似性),一种新的基于GO的基因集相似性度量,可以在线性时间内计算基因集的大小。该度量通过平均每个基因的GO术语及其祖先术语相对于GO词汇表图的贡献来量化一组基因之间的GO注释的相似性。为了研究我们的方法的性能,我们比较了我们的措施与一个既定的基于对的措施时,运行在不同程度的功能相似性的基因集。除了显着的速度提高,我们的方法产生了可比的相似性分数的既定方法。我们的方法可以作为基于Web的工具和开源Python库使用。可用性:基于Web的工具和Python代码可在http://bioserver.cs.rice.edu/gs2上获得。联系人:troy. rice.edu
Motivation: The growing availability of genome-scale datasets has attracted increasing attention to the development of computational methods for automated inference of functional similarities among genes and their products. One class of such methods measures the functional similarity of genes based on their distance in the Gene Ontology (GO). To measure the functional relatedness of a gene set, these measures consider every pair of genes in the set, and the average of all pairwise distances is calculated. However, as more data becomes available and gene sets used for analysis become larger, such pair-based calculation becomes prohibitive. Results: In this article, we propose GS2 (GO-based similarity of gene sets), a novel GO-based measure of gene set similarity that is computable in linear time in the size of the gene set. The measure quantifies the similarity of the GO annotations among a set of genes by averaging the contribution of each gene's GO terms and their ancestor terms with respect to the GO vocabulary graph. To study the performance of our method, we compared our measure with an established pair-based measure when run on gene sets with varying degrees of functional similarities. In addition to a significant speed improvement, our method produced comparable similarity scores to the established method. Our method is available as a web-based tool and an open-source Python library. Availability: The web-based tools and Python code are available at: http://bioserver.cs.rice.edu/gs2. Contact: troy.ruths@rice.edu
DOI: 10.1186/1471-2105-7-470
发表时间: 2006-10-24
期刊: BMC bioinformatics
影响因子: 3
作者:
Beisvag V;Jünge FK;Bergum H;Jølsum L;Lydersen S;Günther CC;Ramampiaro H;Langaas M;Sandvik AK;Laegreid A
通讯作者: Laegreid A
DOI: 10.1093/nar/gkm415
发表时间: 2007-07
影响因子: 14.9
作者:
Huang DW;Sherman BT;Tan Q;Kir J;Liu D;Bryant D;Guo Y;Stephens R;Baseler MW;Lane HC;Lempicki RA
通讯作者: Lempicki RA
DOI: 10.1073/pnas.012025199
发表时间: 2002-04-02
影响因子: 11.1
作者:
Su, AI;Cooke, MP;Hogenesch, JB
通讯作者: Hogenesch, JB
DOI: 10.1109/tcbb.2005.50
发表时间: 2005-10-01
影响因子: 4.5
作者:
Sevilla, JL;Segura, V;Rubio, A
通讯作者: Rubio, A
DOI: 10.1093/bioinformatics/bti551
发表时间: 2005-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Maere, S;Heymans, K;Kuiper, M
通讯作者: Kuiper, M