A large-scale proteogenomics study of apicomplexan pathogens-Toxoplasma gondii and Neospora caninum.

A large-scale proteogenomics study of apicomplexan pathogens-Toxoplasma gondii and Neospora caninum.
复制标题

DOI:
10.1002/pmic.201400553
复制
发表时间:
2015-08
期刊:
影响因子:
3.4
通讯作者:
Jones AR
Jones AR
中科院分区:
生物学3区
文献类型:
--
作者:
Krishna R;Xia D;Sanderson S;Shanmugasundram A;Vermont S;Bernal A;Daniel-Naguib G;Ghali F;Brunk BP;Roos DS;Wastling JM;Jones AR

文献摘要

被引文献

相似文献

蛋白质组学数据可以补充基因组注释工作,例如用于确认基因模型或纠正基因注释错误。在这里,我们提出了一个大规模的蛋白质组学研究的两个重要的顶复门病原体:弓形虫和犬新孢子虫。我们针对直接从RNASeq数据生成的一组官方和替代基因模型查询蛋白质组学数据,使用几个新生成的和一些先前发表的MS数据集进行荟萃分析。我们共鉴定了201 996和39 953个T. gondii和N.犬,分别在1%肽FDR阈值。这相当于鉴定了T. gondii,8911肽/1273蛋白;严格的蛋白质水平阈值后犬。我们还确定了289和140个位点的T。gondii和N. caninum,其映射到我们的分析中使用的RNA-Seq衍生的基因模型,并且显然不存在于这些物种的官方注释(来自EuPathDB的版本10)中。我们在研究中提出了几个例子,其中RNA-Seq证据可以帮助纠正当前的基因模型,并有助于发现潜在的新基因。这项研究的结果已被整合到EuPathDB中。这些数据已存入ProteomeXchange,标识符为PXD 000297和PXD 000298。
Proteomics data can supplement genome annotation efforts, for example being used to confirm gene models or correct gene annotation errors. Here, we present a large-scale proteogenomics study of two important apicomplexan pathogens: Toxoplasma gondii and Neospora caninum. We queried proteomics data against a panel of official and alternate gene models generated directly from RNASeq data, using several newly generated and some previously published MS datasets for this meta-analysis. We identified a total of 201 996 and 39 953 peptide-spectrum matches for T. gondii and N. caninum, respectively, at a 1% peptide FDR threshold. This equated to the identification of 30 494 distinct peptide sequences and 2921 proteins (matches to official gene models) for T. gondii, and 8911 peptides/1273 proteins for N. caninum following stringent protein-level thresholding. We have also identified 289 and 140 loci for T. gondii and N. caninum, respectively, which mapped to RNA-Seq-derived gene models used in our analysis and apparently absent from the official annotation (release 10 from EuPathDB) of these species. We present several examples in our study where the RNA-Seq evidence can help in correction of the current gene model and can help in discovery of potential new genes. The findings of this study have been integrated into the EuPathDB. The data have been deposited to the ProteomeXchange with identifiers PXD000297and PXD000298.