Improvements to the Rice Genome Annotation Through Large-Scale Analysis of RNA-Seq and Proteomics Data Sets.
Improvements to the Rice Genome Annotation Through Large-Scale Analysis of RNA-Seq and Proteomics Data Sets.
复制标题
DOI:
10.1074/mcp.ra118.000832
复制
发表时间:
2019-01
期刊:
影响因子:
--
通讯作者:
Jones AR
中科院分区:
文献类型:
--
作者:
Ren Z;Qi D;Pugh N;Li K;Wen B;Zhou R;Xu S;Liu S;Jones AR
The genome of rice has been sequenced in the past, serving as a starting point to search for useful genetic traits. However, the process of annotating all the genes is a challenging and on-going process. We have re-analyzed large amounts of publicly available data about rice proteins, to correct errors in gene sequences and find new genes. Our new results are presented in a simple format to allow users of databases to see the correspondence between genes and the proteins. Highlights We have mapped public proteomics data against the rice genome/transcriptome. We discovered 1584 novel peptides not currently explained by gene models. 101 new loci were matched by novel peptides, not currently annotated as genes. Data are made persistently available for simple visualization on genome browsers. Rice (Oryza sativa) is one of the most important worldwide crops. The genome has been available for over 10 years and has undergone several rounds of annotation. We created a comprehensive database of transcripts from 29 public RNA sequencing data sets, officially predicted genes from Ensembl plants, and common contaminants in which to search for protein-level evidence. We re-analyzed nine publicly accessible rice proteomics data sets. In total, we identified 420K peptide spectrum matches from 47K peptides and 8,187 protein groups. 4168 peptides were initially classed as putative novel peptides (not matching official genes). Following a strict filtration scheme to rule out other possible explanations, we discovered 1,584 high confidence novel peptides. The novel peptides were clustered into 692 genomic loci where our results suggest annotation improvements. 80% of the novel peptides had an ortholog match in the curated protein sequence set from at least one other plant species. For the peptides clustering in intergenic regions (and thus potentially new genes), 101 loci were identified, for which 43 had a high-confidence hit for a protein domain. Our results can be displayed as tracks on the Ensembl genome or other browsers supporting Track Hubs, to support re-annotation of the rice genome.