Homology-based gene structure prediction: simplified matching algorithm using a translated codon (tron) and improved accuracy by allowing for long gaps

Homology-based gene structure prediction: simplified matching algorithm using a translated codon (tron) and improved accuracy by allowing for long gaps
复制标题

DOI:
10.1093/bioinformatics/16.3.190
复制
发表时间:
2000-03-01
期刊:
影响因子:
5.8
通讯作者:
Gotoh, O
Gotoh, O
中科院分区:
生物学3区
文献类型:
--
作者:
Gotoh, O

文献摘要

被引文献

相似文献

动机:在真核生物基因组DNA序列上定位蛋白质编码外显子(CDS)是预测基因组该部分所包含基因功能的首要且关键的步骤。通过将DNA序列与已知蛋白质序列或同源家族成员的特征图谱直接匹配,或许能够实现对CDS的准确预测。 结果:设计了一种将DNA序列编码为一系列23种可能字母(翻译密码子或tron码)的新规则,以改进此类分析。利用这一规则,开发了一种动态规划算法,用于比对DNA序列和蛋白质序列或特征图谱,使得剪接和翻译后的序列与参考序列的匹配达到最优,就像标准的蛋白质序列比对一样,允许存在长的空位。目标函数还考虑了移码错误、编码潜能以及翻译起始、终止和剪接信号。该方法在已知结构的秀丽隐杆线虫基因上进行了测试。对于所测试的288个基因,以相关系数(CC)衡量的预测准确性在核苷酸水平上约为95%,对于其产物与最相近同源物具有超过30%相同氨基酸的170个基因,准确性为97.0%。我们还提出了一种策略,通过迭代基因预测以及从预测序列中推导出参考特征图谱的重建,来提高一组旁系同源基因的预测准确性。
Motivation: Locating protein-coding exons (CDSs) on a eukaryotic genomic DNA sequence is the initial and an essential step in predicting the functions of the genes embedded in that part of the genome. Accurate prediction of CDSs may be achieved by directly matching the DNA sequence with a known protein sequence or profile of ct homologous family member(s).Results: A new convention for encoding a DNA sequence into a series of 23 possible letters (translated codon or tron code) was devised to improve this type of analysis. Using this convention, a dynamic programming algorithm was developed to align a DNA sequence and a protein sequence or profile so that the spliced and translated sequence optimally matches the reference the same as the standard protein sequence alignment allowing for long gaps. The objective function also takes account of frameshift errors, coding potentials, and translational initiation, termination and splicing signals. This method was tested on Caenorhabditis elegans genes of known structures. The accuracy of prediction measured in terms of a correlation coefficient (CC) was about 95% at the nucleotide level for the 288 genes tested and 97.0% for the 170 genes whose product and closest homologue share more than 30% identical amino acids. We also propose a strategy to improve the accuracy of prediction for a set of paralogous genes by means of iterative gene prediction and reconstruction of the reference profile derived from the predicted sequences.