Using multiple alignments to improve gene prediction

Using multiple alignments to improve gene prediction
复制标题

DOI:
10.1089/cmb.2006.13.379
复制
发表时间:
2006-03-01
影响因子:
1.7
通讯作者:
Brent, MR
Brent, MR
中科院分区:
生物学4区
文献类型:
--
作者:
Gross, SS;Brent, MR

文献摘要

被引文献

相似文献

多物种从头基因预测问题可以表述如下:给定来自两个或多个生物体的基因组序列的比对,预测一个或多个序列中所有蛋白质编码基因的位置和结构。在这里,我们提出了一个新系统 N-SCAN(又名 TWINSCAN 3.0)来解决这个问题。 N-SCAN 可以对比对的基因组序列、上下文相关的替换率以及插入和删除之间的系统发育关系进行建模。创建了 N-SCAN 的实现,并用于生成整个人类基因组和果蝇果蝇基因组的预测。对预测的分析表明,N-SCAN 在人类和果蝇中的准确性超过了之前发布的所有全基因组从头基因预测器。
The multiple species de novo gene prediction problem can be stated as follows: given an alignment of genomic sequences from two or more organisms, predict the location and structure of all protein-coding genes in one or more of the sequences. Here, we present a new system, N-SCAN ( a.k.a. TWINSCAN 3.0), for addressing this problem. N-SCAN can model the phylogenetic relationships between the aligned genome sequences, context-dependent substitution rates, and insertions and deletions. An implementation of N-SCAN was created and used to generate predictions for the entire human genome and the genome of the fruit fly Drosophila melanogaster. Analyses of the predictions reveal that N-SCAN's accuracy in both human and fly exceeds that of all previously published whole-genome de novo gene predictors.