Large-scale RACE approach for proactive experimental definition of C. elegans ORFeome

Large-scale RACE approach for proactive experimental definition of C. elegans ORFeome
复制标题

DOI:
10.1101/gr.098640.109
复制
发表时间:
2009-12-01
期刊:
影响因子:
7
通讯作者:
Vidal, Marc
Vidal, Marc
中科院分区:
生物学1区
文献类型:
--
作者:
Salehi-Ashtiani, Kourosh;Lin, Chenwei;Vidal, Marc

文献摘要

被引文献

相似文献

虽然一个高度精确的秀丽隐杆线虫基因组序列已经存在了10年,但其许多蛋白质编码基因的确切转录结构仍然不确定。大约三分之二的ORFeome已经通过扩增和克隆计算预测的转录模型进行了反应性验证;但仍有整整三分之一的ORFeome未经实验验证。为了充分识别蠕虫基因组的蛋白质编码潜力,包括可能不满足现有基因预测启发式的转录本,我们开发了一个适用于大规模结构转录本注释的cDNA末端快速扩增(RACE)的计算和实验平台。我们使用这个平台询问了2000个未经验证的蛋白质编码基因。我们获得了大约三分之二被检查的转录本的RACE数据,并重建了其中近1000个转录本的ORF和转录本模型。我们定义了未翻译区域,鉴定了新的外显子,并重新定义了先前注释的外显子。我们的结果表明,多达20%的秀丽隐杆线虫基因组可能被错误地注释。我们的大规模RACE平台可以主动纠正许多标注错误。
Although a highly accurate sequence of the Caenorhabditis elegans genome has been available for 10 years, the exact transcript structures of many of its protein-coding genes remain unsettled. Approximately two-thirds of the ORFeome has been verified reactively by amplifying and cloning computationally predicted transcript models; still a full third of the ORFeome remains experimentally unverified. To fully identify the protein-coding potential of the worm genome including transcripts that may not satisfy existing heuristics for gene prediction, we developed a computational and experimental platform adapting rapid amplification of cDNA ends (RACE) for large-scale structural transcript annotation. We interrogated 2000 unverified protein-coding genes using this platform. We obtained RACE data for approximately two-thirds of the examined transcripts and reconstructed ORF and transcript models for close to 1000 of these. We defined untranslated regions, identified new exons, and redefined previously annotated exons. Our results show that as much as 20% of the C. elegans genome may be incorrectly annotated. Many annotation errors could be corrected proactively with our large-scale RACE platform.