Improvements to the Rice Genome Annotation Through Large-Scale Analysis of RNA-Seq and Proteomics Data Sets.

Improvements to the Rice Genome Annotation Through Large-Scale Analysis of RNA-Seq and Proteomics Data Sets.
复制标题

DOI:
10.1074/mcp.ra118.000832
复制
发表时间:
2019-01
期刊:
Molecular & cellular proteomics : MCP
影响因子:
--
通讯作者:
Jones AR
Jones AR
中科院分区:
其他
文献类型:
--
作者:
Ren Z;Qi D;Pugh N;Li K;Wen B;Zhou R;Xu S;Liu S;Jones AR

文献摘要

相似文献

水稻基因组过去已被测序,作为寻找有用遗传性状的起点。然而,注释所有基因的过程是一个具有挑战性且持续的过程。我们重新分析了大量有关水稻蛋白质的公开数据,以纠正基因序列中的错误并找到新的基因。我们的新结果以简单的格式呈现,允许数据库用户查看基因和蛋白质之间的对应关系。亮点 我们已经根据水稻基因组/转录组绘制了公共蛋白质组学数据。我们发现了 1584 种目前无法用基因模型解释的新肽。 101 个新基因座与新肽相匹配,目前尚未注释为基因。数据持续可用,可在基因组浏览器上进行简单可视化。水稻(Oryza sativa)是世界上最重要的农作物之一。该基因组已经面世十多年,并经历了多轮注释。我们创建了一个综合数据库,其中包含 29 个公共 RNA 测序数据集的转录本、来自 Ensembl 植物的正式预测基因以及用于搜索蛋白质水平证据的常见污染物。我们重新分析了九个可公开访问的水稻蛋白质组学数据集。我们总共从 47K 肽和 8,187 个蛋白质组中鉴定出 420K 肽谱匹配。 4168 种肽最初被归类为推定的新型肽(与官方基因不匹配)。经过严格的过滤方案以排除其他可能的解释,我们发现了 1,584 种高置信度的新型肽。新的肽被聚类到 692 个基因组位点,我们的结果表明注释有所改进。 80% 的新型肽在来自至少一种其他植物物种的精选蛋白质序列集中具有直系同源匹配。对于聚集在基因间区域(因此可能是新基因)的肽,鉴定出 101 个基因座,其中 43 个对蛋白质结构域具有高置信度命中。我们的结果可以在 Ensembl 基因组或其他支持 Track Hub 的浏览器上显示为轨迹,以支持水稻基因组的重新注释。
The genome of rice has been sequenced in the past, serving as a starting point to search for useful genetic traits. However, the process of annotating all the genes is a challenging and on-going process. We have re-analyzed large amounts of publicly available data about rice proteins, to correct errors in gene sequences and find new genes. Our new results are presented in a simple format to allow users of databases to see the correspondence between genes and the proteins. Highlights We have mapped public proteomics data against the rice genome/transcriptome. We discovered 1584 novel peptides not currently explained by gene models. 101 new loci were matched by novel peptides, not currently annotated as genes. Data are made persistently available for simple visualization on genome browsers. Rice (Oryza sativa) is one of the most important worldwide crops. The genome has been available for over 10 years and has undergone several rounds of annotation. We created a comprehensive database of transcripts from 29 public RNA sequencing data sets, officially predicted genes from Ensembl plants, and common contaminants in which to search for protein-level evidence. We re-analyzed nine publicly accessible rice proteomics data sets. In total, we identified 420K peptide spectrum matches from 47K peptides and 8,187 protein groups. 4168 peptides were initially classed as putative novel peptides (not matching official genes). Following a strict filtration scheme to rule out other possible explanations, we discovered 1,584 high confidence novel peptides. The novel peptides were clustered into 692 genomic loci where our results suggest annotation improvements. 80% of the novel peptides had an ortholog match in the curated protein sequence set from at least one other plant species. For the peptides clustering in intergenic regions (and thus potentially new genes), 101 loci were identified, for which 43 had a high-confidence hit for a protein domain. Our results can be displayed as tracks on the Ensembl genome or other browsers supporting Track Hubs, to support re-annotation of the rice genome.