Test data sets and evaluation of gene prediction programs on the rice genome

Test data sets and evaluation of gene prediction programs on the rice genome
复制标题

DOI:
10.1007/s11390-005-0446-x
复制
发表时间:
2005-07-01
影响因子:
1.9
通讯作者:
Hao, BL
Hao, BL
中科院分区:
计算机科学3区
文献类型:
--
作者:
Li, H;Liu, JS;Hao, BL

文献摘要

被引文献

相似文献

随着多个水稻基因组计划的临近完成,利用计算机算法进行基因预测/发现已成为一个紧迫的任务。将新发表的28,469个KOME全长水稻cDNA与Oryza sativa ssp的RGP BAC克隆序列进行比对,构建了两个测试集。粳稻:单基因组550个序列,多基因组62个序列,271个基因。这些数据集用于评估5个从头开始的基因预测程序:RiceHMM、GlimmerR、GeneMark、FGENSH和BGF。采用常用的测量方法和几种新的测量方法,在核苷酸、外显子和全基因结构水平上对预测结果进行了比较。测试结果按时间顺序显示成绩的进步。同时,程序的互补性暗示了进一步改进的可能性,以及通过组合几个基因查找器达到更好性能的可行性。
With several rice genome projects approaching completion gene prediction/finding by computer algorithms has become an urgent task. Two test sets were constructed by mapping the newly published 28,469 full-length KOME rice cDNA to the RGP BAC clone sequences of Oryza sativa ssp. japonica: a single-gene set of 550 sequences and a multi-gene set of 62 sequences with 271 genes. These data sets were used to evaluate five ab initio gene prediction programs: RiceHMM, GlimmerR, GeneMark, FGENSH and BGF. The predictions were compared on nucleotide, exon and whole gene structure levels using commonly accepted measures and several new measures. The test results show a progress in performance in chronological order. At the same time complementarity of the programs hints on the possibility of further improvement and on the feasibility of reaching better performance by combining several gene-finders.