Discovery of novel genes and gene isoforms by integrating transcriptomic and proteomic profiling from mouse liver.

Discovery of novel genes and gene isoforms by integrating transcriptomic and proteomic profiling from mouse liver.
复制标题

DOI:
10.1021/pr4012206
复制
发表时间:
2014-04
影响因子:
4.4
通讯作者:
Peng Wu;Hongyu Zhang;Weiran Lin;Yunwei Hao;L. Ren;Chengpu Zhang;Ning Li;Handong Wei;Ying Jiang;F. He
Peng Wu;Hongyu Zhang;Weiran Lin;Yunwei Hao;L. Ren;Chengpu Zhang;Ning Li;Handong Wei;Ying Jiang;F. He
中科院分区:
生物学2区
文献类型:
--
作者:
Peng Wu;Hongyu Zhang;Weiran Lin;Yunwei Hao;L. Ren;Chengpu Zhang;Ning Li;Handong Wei;Ying Jiang;F. He

文献摘要

被引文献

相似文献

在转录组和蛋白质组水平上全面鉴定一种组织的基因表达是深入了解其生物学功能的先决条件。选择性剪接和RNA编辑是转录加工的两种主要形式,在转录组和蛋白质组多样性中起着重要作用,导致同一基因存在多个异构体,而标准蛋白质数据库中异构体信息相对缺乏,难以通过质谱(MS)蛋白质组学方法鉴定。在我们的研究中,我们将MS和RNA-Seq平行用于小鼠肝脏组织,并捕获了相当多的转录本和蛋白质,分别覆盖了Ensembl中60%和34%的蛋白质编码基因。然后,我们开发了一个生物信息学工作流程,用于构建一个定制的蛋白质数据库,该数据库首次包括新的剪接衍生肽和RNA编辑引起的肽变体,使我们能够更完整地识别蛋白质亚型。使用这个实验确定的数据库,我们总共鉴定了150个在标准生物数据库中不存在的肽,错误发现率<1%,对应于72个新的剪接异构体,43个新的遗传区域和15个RNA编辑位点。其中,11个随机选择的新事件通过了PCR和桑格测序的实验验证。新发现的基因产物在两个组学水平上具有高置信度,证明了我们的方法的鲁棒性和有效性及其在改进基因组注释中的潜在应用。所有MS数据均已存入iProx(http://ww.iprox.org),标识符为IPX 00003601。
Comprehensively identifying gene expression in both transcriptomic and proteomic levels of one tissue is a prerequisite for a deeper understanding of its biological functions. Alternative splicing and RNA editing, two main forms of transcriptional processing, play important roles in transcriptome and proteome diversity and result in multiple isoforms for one gene, which are hard to identify by mass spectrometry (MS)-based proteomics approach due to the relative lack of isoform information in standard protein databases. In our study, we employed MS and RNA-Seq in parallel into mouse liver tissue and captured a considerable catalogue of both transcripts and proteins that, respectively, covered 60 and 34% of protein-coding genes in Ensembl. We then developed a bioinformatics workflow for building a customized protein database that for the first time included new splicing-derived peptides and RNA-editing-caused peptide variants, allowing us to more completely identify protein isoforms. Using this experimentally determined database, we totally identified 150 peptides not present in standard biological databases at false discovery rate of <1%, corresponding to 72 novel splicing isoforms, 43 new genetic regions, and 15 RNA-editing sites. Of these, 11 randomly selected novel events passed experimental verification by PCR and Sanger sequencing. New discoveries of gene products with high confidence in two omics levels demonstrated the robustness and effectiveness of our approach and its potential application into improve genome annotation. All the MS data have been deposited to the iProx ( http://ww.iprox.org ) with the identifier IPX00003601.