Exploiting proteomic data for genome annotation and gene model validation in Aspergillus niger.

Exploiting proteomic data for genome annotation and gene model validation in Aspergillus niger.
复制标题

DOI:
10.1186/1471-2164-10-61
复制
发表时间:
2009-02-04
期刊:
影响因子:
4.4
通讯作者:
Hubbard SJ
Hubbard SJ
中科院分区:
生物学2区
文献类型:
--
作者:
Wright JC;Sugden D;Francis-McIntyre S;Riba-Garcia I;Gaskell SJ;Grigoriev IV;Baker SE;Beynon RJ;Hubbard SJ

文献摘要

参考文献

被引文献

相似文献

蛋白质组学数据是一个潜在的丰富的,但有争议的未开发的基因组注释数据源。肽鉴定从串联质谱提供初步证据的基因预测和可以区分一组候选基因模型。在这里,我们将此应用于最近测序的来自联合基因组研究所(JGI)的黑曲霉真菌基因组和来自另一个黑曲霉序列的另一个预测蛋白集。从1d凝胶电泳带获得串联质谱(MS/MS),并使用平均肽评分(APS)和反向数据库搜索对所有可用的基因模型进行搜索,以在可接受的错误发现率(FDR)下产生可靠的鉴定。405个鉴定的肽序列与214个不同的黑曲霉基因组位点相关联,4093个预测基因模型聚集在这些位点上,其中2872个含有所鉴定的肽。有趣的是,这些位点中有13个(6%)没有首选的预测基因模型,或者基因组注释者为该基因组位点选择的“最佳”模型没有被发现与鉴定的肽最简洁匹配。所鉴定的多肽也提高了对来自不同基因模型的54个内含子的预测基因结构的信心。这项工作强调了将实验蛋白质组学数据整合到基因组注释管道中的潜力,就像表达序列标签(EST)数据一样。与DSM测序的另一株黑曲霉基因组的比较表明,许多具有蛋白质组学证据的基因模型或蛋白质在两个基因组中都没有出现,进一步突出了该方法的实用性。
Proteomic data is a potentially rich, but arguably unexploited, data source for genome annotation. Peptide identifications from tandem mass spectrometry provide prima facie evidence for gene predictions and can discriminate over a set of candidate gene models. Here we apply this to the recently sequenced Aspergillus niger fungal genome from the Joint Genome Institutes (JGI) and another predicted protein set from another A.niger sequence. Tandem mass spectra (MS/MS) were acquired from 1d gel electrophoresis bands and searched against all available gene models using Average Peptide Scoring (APS) and reverse database searching to produce confident identifications at an acceptable false discovery rate (FDR). 405 identified peptide sequences were mapped to 214 different A.niger genomic loci to which 4093 predicted gene models clustered, 2872 of which contained the mapped peptides. Interestingly, 13 (6%) of these loci either had no preferred predicted gene model or the genome annotators' chosen "best" model for that genomic locus was not found to be the most parsimonious match to the identified peptides. The peptides identified also boosted confidence in predicted gene structures spanning 54 introns from different gene models. This work highlights the potential of integrating experimental proteomics data into genomic annotation pipelines much as expressed sequence tag (EST) data has been. A comparison of the published genome from another strain of A.niger sequenced by DSM showed that a number of the gene models or proteins with proteomics evidence did not occur in both genomes, further highlighting the utility of the method.
DOI: 10.1186/1471-2164-8-255
发表时间: 2007-07-27
期刊: BMC genomics
影响因子: 4.4
作者:
Lu F;Jiang H;Ding J;Mu J;Valenzuela JG;Ribeiro JM;Su XZ
通讯作者: Su XZ
DOI: 10.1002/pmic.200500648
发表时间: 2006-05-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
McCarthy, Fiona M.;Cooksey, Amanda M.;Burgess, Shane C.
通讯作者: Burgess, Shane C.
DOI: 10.1101/gr.10.4.547
发表时间: 2000-04-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Birney, E;Durbin, R
通讯作者: Durbin, R
DOI: 10.1038/nbt1282
发表时间: 2007-02-01
影响因子: 46.9
作者:
Pel, Herman J.;de Winde, Johannes H.;Stam, Hein
通讯作者: Stam, Hein
DOI: 10.1186/gb-2006-7-4-r35
发表时间: 2006
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Fermin, Damian;Allen, Baxter B.;Blackwell, Thomas W.;Menon, Rajasree;Adamski, Marcin;Xu, Yin;Ulintz, Peter;Omenn, Gilbert S.;States, David J.
通讯作者: States, David J.