Improvements to the rice genome annotation through large-scale analysis of RNA-Seq and proteomics datasets

Improvements to the rice genome annotation through large-scale analysis of RNA-Seq and proteomics datasets
复制标题

通过大规模 RNA-Seq 和蛋白质组学数据集分析改进水稻基因组注释

DOI:
10.1101/300426
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
Ren Z
Ren Z
中科院分区:
--
文献类型:
--
作者:
Ren Z

文献摘要

相似文献

水稻(Oryza sativa)是世界上最重要的农作物之一。该基因组已经存在了10多年,并经历了几轮注释。我们从29个公共RNA测序数据集创建了一个全面的转录本数据库,正式预测了Ensembl植物的基因,以及常见的污染物,以寻找蛋白质水平的证据。我们重新分析了9个公开的水稻蛋白质组学数据集。总共,我们从47 K肽和8,187个蛋白质组中鉴定出420 K肽谱匹配。4168肽最初被归类为推定的新肽(不匹配官方基因)。经过严格的筛选,排除了其他可能的解释,我们发现了1,584个高置信度的新肽。新的肽被聚类到692个基因组位点,我们的结果表明注释改进。80%的新肽在来自至少一种其他植物物种的精选蛋白质序列集中具有直系同源物匹配。对于聚集在基因间区域的肽(因此可能是新的基因),鉴定了101个位点,其中43个具有高置信度的蛋白质结构域命中。我们的结果可以在Ensembl基因组或其他支持Track Hubs的浏览器上显示为轨迹,以支持水稻基因组的重新注释。
Rice (Oryza sativa) is one of the most important worldwide crops. The genome has been available for over 10 years and has undergone several rounds of annotation. We created a comprehensive database of transcripts from 29 public RNA sequencing data sets, officially predicted genes from Ensembl plants, and common contaminants in which to search for protein-level evidence. We re-analyzed nine publicly accessible rice proteomics data sets. In total, we identified 420K peptide spectrum matches from 47K peptides and 8,187 protein groups. 4168 peptides were initially classed as putative novel peptides (not matching official genes). Following a strict filtration scheme to rule out other possible explanations, we discovered 1,584 high confidence novel peptides. The novel peptides were clustered into 692 genomic loci where our results suggest annotation improvements. 80% of the novel peptides had an ortholog match in the curated protein sequence set from at least one other plant species. For the peptides clustering in intergenic regions (and thus potentially new genes), 101 loci were identified, for which 43 had a high-confidence hit for a protein domain. Our results can be displayed as tracks on the Ensembl genome or other browsers supporting Track Hubs, to support re-annotation of the rice genome.