The KA/KS ratio test for assessing the protein-coding potential of genomic regions:: An empirical and simulation study

The KA/KS ratio test for assessing the protein-coding potential of genomic regions:: An empirical and simulation study
复制标题

DOI:
10.1101/gr.200901
复制
发表时间:
2002-01-01
期刊:
影响因子:
7
通讯作者:
Li, WH
Li, WH
中科院分区:
生物学1区
文献类型:
--
作者:
Nekrutenko, A;Makova, KD;Li, WH

文献摘要

被引文献

相似文献

比较基因组学是提高基因预测准确性的一种简单而有力的方法。在这项研究中,我们展示了一种通过人/鼠序列比较来识别蛋白质编码外显子的简单测试的实用性。该测试利用了在绝大多数编码区中,同义替换(K-S)比非同义替换(K-A)发生得更频繁这一事实,并以K-A/K-S比率作为判断标准。我们发现:(1)大多数人和小鼠的外显子足够长,并且具有合适的序列差异程度,以便可靠地进行测试;(2)该测试适合于识别用现有方法难以预测的长外显子和单外显子基因;(3)该测试的假阴性率低于大多数现有的基因预测方法,假阳性率低于所有现有的方法;(4)该测试已经自动化,并可与其他现有的基因预测方法结合使用。
Comparative genomics is a simple, powerful way to increase the accuracy of gene prediction. In this study, we show the utility of a simple test for the identification of protein-coding exons using human/mouse sequence comparisons. The test takes advantage of the fact that in the vast majority of coding regions, synonymous substitutions (K-S) Occur much more frequently than nonsynonymous ones (K-A) and uses the K-A/K-S ratio as the criterion. We show the following: (1) most of the human and mouse exons are sufficiently long and have a suitable degree of sequence divergence for the test to perform reliably; (2) the test is suited for the identification of long exons and single exon genes, which are difficult to predict by Current methods; (3) the test has a false-negative rate, lower than most Of Current gene prediction methods and a false-positive rate lower than all Current methods; (4) the test has been automated and call be used in combination with other existing gene-prediction methods.