Assessment and refinement of eukaryotic gene structure prediction with gene-structure-aware multiple protein sequence alignment.

Assessment and refinement of eukaryotic gene structure prediction with gene-structure-aware multiple protein sequence alignment.
复制标题

DOI:
10.1186/1471-2105-15-189
复制
发表时间:
2014-06-14
期刊:
影响因子:
3
通讯作者:
Nelson DR
Nelson DR
中科院分区:
生物学4区
文献类型:
--
作者:
Gotoh O;Morita M;Nelson DR

文献摘要

参考文献

被引文献

相似文献

真核生物基因组织的精确计算鉴定是一个长期存在的问题。尽管对新测序基因组中编码的基因进行精确注释具有重要意义,但预测基因结构的准确性尚未得到严格评估,这主要是由于缺乏适当的评估方法。我们提出了一种基因结构敏感的多序列比对方法,用于利用从许多基因组中同源基因翻译的氨基酸序列进行基因预测。该方法提供了有关每个预测基因结构可靠性的丰富信息。我们还设计了一种迭代方法,试图改进基于拼接比对算法的可疑预测基因的结构,使用共识序列或可靠的同源物作为模板。应用我们的方法对47个植物基因组的细胞色素P450和核糖体蛋白进行分析表明,50 ~ 60%的注释基因结构可能含有一些缺陷。然而,超过一半的含有缺陷的基因可能是内在断裂的,即它们是假基因或基因片段,位于未完成的测序区域,或对应于非生产性同工型,而在大多数剩余的候选基因中发现的缺陷可以通过我们的迭代改进方法进行修复。通过基因结构感知多蛋白序列比对介导的真核生物基因结构精化是显著提高一组同源基因整体预测质量的有效策略。我们的方法将适用于各种蛋白质编码基因家族,如果它们的结构域结构是进化稳定的。将我们的方法应用于所有生命王国的基因家族也是可行的,而不仅仅是植物。
Accurate computational identification of eukaryotic gene organization is a long-standing problem. Despite the fundamental importance of precise annotation of genes encoded in newly sequenced genomes, the accuracy of predicted gene structures has not been critically evaluated, mostly due to the scarcity of proper assessment methods. We present a gene-structure-aware multiple sequence alignment method for gene prediction using amino acid sequences translated from homologous genes from many genomes. The approach provides rich information concerning the reliability of each predicted gene structure. We have also devised an iterative method that attempts to improve the structures of suspiciously predicted genes based on a spliced alignment algorithm using consensus sequences or reliable homologs as templates. Application of our methods to cytochrome P450 and ribosomal proteins from 47 plant genomes indicated that 50 ~ 60 % of the annotated gene structures are likely to contain some defects. Whereas more than half of the defect-containing genes may be intrinsically broken, i.e. they are pseudogenes or gene fragments, located in unfinished sequencing areas, or corresponding to non-productive isoforms, the defects found in a majority of the remaining gene candidates can be remedied by our iterative refinement method. Refinement of eukaryotic gene structures mediated by gene-structure-aware multiple protein sequence alignment is a useful strategy to dramatically improve the overall prediction quality of a set of homologous genes. Our method will be applicable to various families of protein-coding genes if their domain structures are evolutionarily stable. It is also feasible to apply our method to gene families from all kingdoms of life, not just plants.
DOI: 10.1093/bioinformatics/8.3.275
发表时间: 1992-06-01
期刊: COMPUTER APPLICATIONS IN THE BIOSCIENCES
影响因子: --
作者:
JONES, DT;TAYLOR, WR;THORNTON, JM
通讯作者: THORNTON, JM
DOI: 10.1093/nar/gkt1183
发表时间: 2014-01
影响因子: 14.9
作者:
Grigoriev IV;Nikitin R;Haridas S;Kuo A;Ohm R;Otillar R;Riley R;Salamov A;Zhao X;Korzeniewski F;Smirnova T;Nordberg H;Dubchak I;Shabalov I
通讯作者: Shabalov I
DOI: 10.1093/bioinformatics/16.3.190
发表时间: 2000-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Gotoh, O
通讯作者: Gotoh, O
DOI: 10.1101/gr.1858004
发表时间: 2004-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Curwen, V;Eyras, E;Clamp, M
通讯作者: Clamp, M
将cDNA序列映射到基因组序列上的空间有效和准确的方法。
DOI: 10.1093/nar/gkn105
发表时间: 2008-05
影响因子: 14.9
作者:
Gotoh, Osamu
通讯作者: Gotoh, Osamu