Annotation of the Zebrafish Genome through an Integrated Transcriptomic and Proteomic Analysis

Annotation of the Zebrafish Genome through an Integrated Transcriptomic and Proteomic Analysis
复制标题

DOI:
10.1074/mcp.m114.038299
复制
发表时间:
2014-11-01
影响因子:
7
通讯作者:
Pandey, Akhilesh
Pandey, Akhilesh
中科院分区:
生物学1区
文献类型:
--
作者:
Kelkar, Dhanashree S.;Provost, Elayne;Pandey, Akhilesh

文献摘要

被引文献

相似文献

蛋白质编码基因的准确注释是完成任何生物体全基因组测序的首要任务之一。在这项研究中,我们使用了整合的转录和蛋白质组学策略来验证和改进现有的斑马鱼基因组注释。除了6个器官的转录图谱外,我们还对斑马鱼的10个成年器官、整个成年鱼身体和两个发育阶段的斑马鱼(SAT品系)进行了高分辨率质谱学蛋白质组图谱的研究。从蛋白质组学分析中鉴定出超过7,000种蛋白质,并从RNA测序数据中组装了大约69,000个高置信度转录本。大约15%的转录本定位于基因间隔区,其中大部分可能是长的非编码RNA。这些高质量的转录和蛋白质组学数据被用来手动重新注释斑马鱼基因组。我们报告了157个新的蛋白质编码基因的鉴定。此外,我们的数据还导致了现有基因结构的修改,包括新的外显子、外显子坐标的变化、翻译框架的变化、注释UTRs中的翻译以及基因的连接。最后,我们发现了四个基因组组装错误的实例,蛋白质组和转录组数据都支持这些错误。我们的研究表明,对转录组和蛋白质组的综合分析可以扩展我们对即使是注释良好的基因组的理解。
Accurate annotation of protein-coding genes is one of the primary tasks upon the completion of whole genome sequencing of any organism. In this study, we used an integrated transcriptomic and proteomic strategy to validate and improve the existing zebrafish genome annotation. We undertook high-resolution mass-spectrometry-based proteomic profiling of 10 adult organs, whole adult fish body, and two developmental stages of zebrafish (SAT line), in addition to transcriptomic profiling of six organs. More than 7,000 proteins were identified from proteomic analyses, and approximate to 69,000 high-confidence transcripts were assembled from the RNA sequencing data. Approximately 15% of the transcripts mapped to intergenic regions, the majority of which are likely long non-coding RNAs. These high-quality transcriptomic and proteomic data were used to manually reannotate the zebrafish genome. We report the identification of 157 novel protein-coding genes. In addition, our data led to modification of existing gene structures including novel exons, changes in exon coordinates, changes in frame of translation, translation in annotated UTRs, and joining of genes. Finally, we discovered four instances of genome assembly errors that were supported by both proteomic and transcriptomic data. Our study shows how an integrative analysis of the transcriptome and the proteome can extend our understanding of even well-annotated genomes.