Computational evaluation of TIS annotation for prokaryotic genomes.

Computational evaluation of TIS annotation for prokaryotic genomes.
复制标题

原核生物基因组 TIS 注释的计算评估

DOI:
10.1186/1471-2105-9-160
复制
发表时间:
2008-03-25
期刊:
影响因子:
3
通讯作者:
She, Zhen-Su
She, Zhen-Su
中科院分区:
生物学4区
文献类型:
--
作者:
Hu, Gang-Qing;Zheng, Xiaobin;Ju, Li-Ning;Zhu, Huaiqiu;She, Zhen-Su

文献摘要

参考文献

被引文献

相似文献

背景翻译起始位点(TIS)的准确注释对于理解翻译起始机制至关重要。然而,由于缺乏实验基准,RefSeq等广泛使用的数据库中TIS注释的可靠性并不确定。结果基于基因翻译相关信号均匀分布在基因组中的同质性假设,我们建立了一种计算方法,用于大规模定量评估任何原核生物基因组的TIS注释的可靠性。该方法包括根据三个基本 PWM 的线性组合对预测 TIS 周围对齐序列的位置权重矩阵 (PWM) 进行建模,其中一个用于真实 TIS,另外两个用于假 TIS。三个基本 PWM 是使用具有高度可靠的 TIS 预测的参考集获得的。广义最小二乘估计器确定观察到的 PWM 中真实 TIS 的权重,从中得出预测的准确性。该方法的有效性和假设的限制程度通过对参考集的可变精度的实验验证的 TIS 进行测试来明确解决。该方法用于估计 RefSeq 和 ProTISA 等公共数据库以及 EasyGene、GeneMarkS、Glimmer 3 和 TiCo 等程序提供的 TIS 注释的准确性。结果表明,RefSeq 的 TIS 预测准确度明显低于最近的两个预测器 Tico 和 ProTISA。通过令人信服的证据,我们在 RefSeq 注释中展示了两种普遍的偏好偏差,即过度注释最长开放阅读框 (LORF) 和不足注释 ATG 起始密码子。最后,我们基于所有预测变量的最佳预测,建立了一个新的TIS数据库SupTISA; SupTISA在所有532个完整基因组上实现了92%的平均准确率。结论已经实现了TIS注释的大规模计算评估。一个比RefSeq更好的新TIS数据库已经建立,它为进一步的TIS研究提供了宝贵的资源。
BackgroundAccurate annotation of translation initiation sites (TISs) is essential for understanding the translation initiation mechanism. However, the reliability of TIS annotation in widely used databases such as RefSeq is uncertain due to the lack of experimental benchmarks.ResultsBased on a homogeneity assumption that gene translation-related signals are uniformly distributed across a genome, we have established a computational method for a large-scale quantitative assessment of the reliability of TIS annotations for any prokaryotic genome. The method consists of modeling a positional weight matrix (PWM) of aligned sequences around predicted TISs in terms of a linear combination of three elementary PWMs, one for true TIS and the two others for false TISs. The three elementary PWMs are obtained using a reference set with highly reliable TIS predictions. A generalized least square estimator determines the weighting of the true TIS in the observed PWM, from which the accuracy of the prediction is derived. The validity of the method and the extent of the limitation of the assumptions are explicitly addressed by testing on experimentally verified TISs with variable accuracy of the reference sets. The method is applied to estimate the accuracy of TIS annotations that are provided on public databases such as RefSeq and ProTISA and by programs such as EasyGene, GeneMarkS, Glimmer 3 and TiCo. It is shown that RefSeq's TIS prediction is significantly less accurate than two recent predictors, Tico and ProTISA. With convincing proofs, we show two general preferential biases in the RefSeq annotation,i.e. over-annotating the longest open reading frame (LORF) and under-annotating ATG start codon. Finally, we have established a new TIS database, SupTISA, based on the best prediction of all the predictors; SupTISA has achieved an average accuracy of 92% over all 532 complete genomes.ConclusionLarge-scale computational evaluation of TIS annotation has been achieved. A new TIS database much better than RefSeq has been constructed, and it provides a valuable resource for further TIS studies.
DOI: 10.1186/1471-2105-4-21
发表时间: 2003-06-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Larsen TS;Krogh A
通讯作者: Krogh A
MED:一种新的细菌和古细菌基因组无监督基因预测算法。
DOI: 10.1186/1471-2105-8-97
发表时间: 2007-03-16
期刊: BMC bioinformatics
影响因子: 3
作者:
Zhu H;Hu GQ;Yang YF;Wang J;She ZS
通讯作者: She ZS
DOI: 10.1073/pnas.71.4.1342
发表时间: 1974-01-01
影响因子: 11.1
作者:
SHINE J;DALGARNO L
通讯作者: DALGARNO L
DOI: 10.1093/nar/29.12.2607
发表时间: 2001-06-15
影响因子: 14.9
作者:
Besemer, J;Lomsadze, A;Borodovsky, M
通讯作者: Borodovsky, M
DOI: 10.1016/s0378-1119(99)00200-0
发表时间: 1999-07-08
期刊: GENE
影响因子: 3.5
作者:
Frishman, D;Mironov, A;Gelfand, M
通讯作者: Gelfand, M