Identifying secretomes in people, pufferfish and pigs

Identifying secretomes in people, pufferfish and pigs
复制标题

DOI:
10.1093/nar/gkh286
复制
发表时间:
2004-02-01
影响因子:
14.9
通讯作者:
Ellis, LBM
Ellis, LBM
中科院分区:
生物学2区
文献类型:
--
作者:
Klee, EW;Carlson, DF;Ellis, LBM

文献摘要

被引文献

相似文献

通过分泌途径(分泌组)加工的蛋白质在多细胞真核生物的发育中起着关键作用,但尚未在基因组水平上进行全面研究。在这项研究中,我们使用Target P算法基于对人类(13-20%在单个数据集中发现的蛋白质)和河豚(14%)几乎完整的蛋白质组的分析来预测人类(13-20%在单个数据集中发现的蛋白质)和河豚(14%)的分泌组。我们将内部处理与预测软件相结合,以自动识别分泌蛋白,并克服与EST数据相关的主要挑战之一:识别编码n端完全蛋白的少数克隆。我们讨论了使用这些方法来预测基于est的共识序列集的分泌蛋白,并使用无细胞共翻译易位分析验证了这些预测。以TIGR猪基因指数4.0作为测试数据集进行分析,鉴定出352个n端完整的推定分泌蛋白。在功能上与我们的预测一致,40个(85%)这些cdna中有34个被证实在体外翻译系统中共翻译易位。本文开发的方法专门用于接受部分开放阅读框,并改善真核生物转录组中分泌蛋白的预测,对真核生物EST数据库的分析和注释有价值。
The proteins processed by the secretory pathway (secretome) are critical players in the development of multi-cellular eukaryotic organisms but have yet to be comprehensively studied at the genomic level. In this study, we use the Target P algorithm to predict human (13-20% of proteins found in individual datasets) and Fugu (14%) secretomes based on analysis of their nearly complete proteomes. We combine internal processing with prediction software to automate secreted protein identification and overcome one of the major challenges associated with EST data: identification of the minority of clones that encode N-terminally-complete proteins. We discuss the use of these methods to predict secreted proteins in EST-based consensus sequence sets, and we validate these predictions using an assay for cell-free cotranslational translocation. Analysis of TIGR Porcine Gene Index 4.0 as a test dataset resulted in the identification of 352 N-terminally-complete, putative secreted proteins. In functional agreement with our predictions, 34 of 40 (85%) of these cDNAs were verified to be cotranslationally translocated in an in vitro translation system. The methods developed here are specifically designed to accept partial open reading frames and improve secreted protein predictions in eukaryotic transcriptomes, and are valuable for the analysis and annotation of eukaryotic EST databases.