Experimental and statistical post-validation of positive example EST sequences carrying peroxisome targeting signals type 1 (PTS1)

Experimental and statistical post-validation of positive example EST sequences carrying peroxisome targeting signals type 1 (PTS1)
复制标题

携带过氧化物酶体靶向信号 1 型 (PTS1) 的阳性 EST 序列的实验和统计后验证

DOI:
10.4161/psb.18720
复制
发表时间:
2012
影响因子:
2.9
通讯作者:
S. Reumann
S. Reumann
中科院分区:
生物学4区
文献类型:
--
作者:
Thomas Lingner;Amr R. A. Kataya;S. Reumann

文献摘要

被引文献

相似文献

我们最近开发了第一个专门针对植物的算法,用于从基因组序列中预测携带过氧化物酶体靶向信号 1 型 (PTS1) 的蛋白质。1经实验验证,预测方法能够正确预测未知的过氧化物酶体拟南芥蛋白质并推断新的 PTS1 三肽。高预测性能主要取决于底层正例序列的数量大和序列多样性,这些序列主要来源于EST数据库。然而,在实验验证研究中,一些构建体仍保留在胞质中,这表明某些 EST 中存在测序错误。为了识别错误序列,我们在本研究中验证了其他阳性示例序列的亚细胞靶向。此外,我们分别分析了 PTS1 蛋白的每个直系同源组的预测分数分布,其通常类似于具有特定组平均值的正态分布。胞质序列通常代表低预测分数的异常值,并且位于拟合正态分布的最尾部。对三种识别异常值的统计方法的敏感性和特异性进行了比较。”它们的组合应用可以从正例数据集中消除错误的 EST。这种新的后验证方法将进一步提高植物、真菌和哺乳动物的 PTS1 和 PTS2 蛋白质预测模型的预测准确性。
We recently developed the first algorithms specifically for plants to predict proteins carrying peroxisome targeting signals type 1 (PTS1) from genome sequences.1 As validated experimentally, the prediction methods are able to correctly predict unknown peroxisomal Arabidopsis proteins and to infer novel PTS1 tripeptides. The high prediction performance is primarily determined by the large number and sequence diversity of the underlying positive example sequences, which mainly derived from EST databases. However, a few constructs remained cytosolic in experimental validation studies, indicating sequencing errors in some ESTs. To identify erroneous sequences, we validated subcellular targeting of additional positive example sequences in the present study. Moreover, we analyzed the distribution of prediction scores separately for each orthologous group of PTS1 proteins, which generally resembled normal distributions with group-specific mean values. The cytosolic sequences commonly represented outliers of low prediction scores and were located at the very tail of a fitted normal distribution. Three statistical methods for identifying outliers were compared in terms of sensitivity and specificity.” Their combined application allows elimination of erroneous ESTs from positive example data sets. This new post-validation method will further improve the prediction accuracy of both PTS1 and PTS2 protein prediction models for plants, fungi, and mammals.