Benchmarking ortholog identification methods using functional genomics data.

Benchmarking ortholog identification methods using functional genomics data.
复制标题

使用功能基因组学数据的基准测试直系同源识别方法。

DOI:
10.1186/gb-2006-7-4-r31
复制
发表时间:
2006
期刊:
影响因子:
12.3
通讯作者:
Groenen PM
Groenen PM
中科院分区:
生物学1区
文献类型:
--
作者:
Hulsen T;Huynen MA;de Vlieg J;Groenen PM

文献摘要

参考文献

被引文献

相似文献

使用功能基因组学数据对最流行的正基因识别方法进行基准测试,确定了两种最佳方法。从模式生物蛋白质到人类蛋白质的功能注释转移是比较基因组学的主要应用之一。各种方法被用来分析跨物种的直系关系,根据一个操作定义的直系关系。通常,直系同源的定义被错误地解释为预测不同物种之间功能等同的蛋白质,而实际上它只定义了不同物种中基因的共同祖先的存在。然而,已经证明,直系同源物经常显示出显著的功能相似性。因此,正字法预测的质量是传递功能注释(和其他相关信息)的重要因素。为了鉴定具有尽可能高的功能相似性的蛋白质对,重要的是要鉴定直系同源物鉴定方法。为了测量来自不同物种的蛋白质功能的相似性,我们使用功能基因组学数据,例如表达数据和蛋白质相互作用数据。我们测试了几种最流行的直系同源物鉴定方法。在一般情况下,我们观察到一个敏感性/选择性的权衡:功能相似性得分每orthophosphorus对序列变得更高时,包括在直向同源物组中的蛋白质的数量减少。通过将灵敏度和选择性结合到总体得分中,我们表明InParanoid程序在识别功能等效蛋白方面是最好的直系同源物识别方法。
A benchmarking of the most popular orthologous identification methods using functional genomics data identifies the two best methods. The transfer of functional annotations from model organism proteins to human proteins is one of the main applications of comparative genomics. Various methods are used to analyze cross-species orthologous relationships according to an operational definition of orthology. Often the definition of orthology is incorrectly interpreted as a prediction of proteins that are functionally equivalent across species, while in fact it only defines the existence of a common ancestor for a gene in different species. However, it has been demonstrated that orthologs often reveal significant functional similarity. Therefore, the quality of the orthology prediction is an important factor in the transfer of functional annotations (and other related information). To identify protein pairs with the highest possible functional similarity, it is important to qualify ortholog identification methods. To measure the similarity in function of proteins from different species we used functional genomics data, such as expression data and protein interaction data. We tested several of the most popular ortholog identification methods. In general, we observed a sensitivity/selectivity trade-off: the functional similarity scores per orthologous pair of sequences become higher when the number of proteins included in the ortholog groups decreases. By combining the sensitivity and the selectivity into an overall score, we show that the InParanoid program is the best ortholog identification method in terms of identifying functionally equivalent proteins.
DOI: 10.1073/pnas.0402591101
发表时间: 2004-06-15
影响因子: 11.1
作者:
Fraser, HB;Hirsh, AE;Eisen, MB
通讯作者: Eisen, MB
DOI: 10.1093/bioinformatics/bth021
发表时间: 2004-01-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Sjölander, K
通讯作者: Sjölander, K
DOI: 10.1093/nar/gkh036
发表时间: 2004-01-01
影响因子: 14.9
作者:
Harris, MA;Clark, J;White, R
通讯作者: White, R
DOI: 10.1001/jama.243.8.756
发表时间: 1980-01-01
影响因子: 120.7
作者:
COTE, RA;ROBBOY, S
通讯作者: ROBBOY, S
DOI: 10.1101/gr.1858004
发表时间: 2004-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Curwen, V;Eyras, E;Clamp, M
通讯作者: Clamp, M