A multi-model statistical approach for proteomic spectral count quantitation.

A multi-model statistical approach for proteomic spectral count quantitation.
复制标题

DOI:
10.1016/j.jprot.2016.05.032
复制
发表时间:
2016-07-20
影响因子:
3.3
通讯作者:
Freitas MA
Freitas MA
中科院分区:
生物学2区
文献类型:
--
作者:
Branson OE;Freitas MA

文献摘要

被引文献

相似文献

随着质谱技术的快速发展,鸟枪法蛋白质组学已成为大规模蛋白质组研究的最有力的分析平台。能够绘制和确定整个蛋白质组的差异表达谱是鸟枪法蛋白质组学的最终目标。无标记定量已被证明是一种有效的方法,发现鸟枪蛋白质组学,特别是当样品是有限的。无标记光谱计数定量是一种类似于RNA测序的方法,其中计数数据用于确定差异表达。在这里,我们表明,统计方法开发,以评估差异表达的RNA测序实验中,可以应用于检测差异蛋白质表达的无标记发现蛋白质组学。这种称为MultiSpec的方法利用开源统计平台,即edgeR,DESeq和baySeq,以统计学方式选择候选蛋白进行进一步研究。此外,为了去除与单一统计方法相关的偏差,通过将edgeR和DESeq q q值与由baySeq计算的错误发现率(FDR)直接比较来组装差异表达蛋白的单一排序列表。这种统计方法,然后扩展时,适用于光谱计数数据来自多个蛋白质组管道。通过折叠蛋白质组的方式整合并交叉验证来自多个蛋白质组学管道的单个统计结果。
The rapid development of mass spectrometry (MS) technologies has solidified shotgun proteomics as the most powerful analytical platform for large-scale proteome interrogation. The ability to map and determine differential expression profiles of the entire proteome is the ultimate goal of shotgun proteomics. Label-free quantitation has proven to be a valid approach for discovery shotgun proteomics, especially when sample is limited. Label-free spectral count quantitation is an approach analogous to RNA sequencing whereby count data is used to determine differential expression. Here we show that statistical approaches developed to evaluate differential expression in RNA sequencing experiments can be applied to detect differential protein expression in label-free discovery proteomics. This approach, termed MultiSpec, utilizes open-source statistical platforms; namely edgeR, DESeq and baySeq, to statistically select protein candidates for further investigation. Furthermore, to remove bias associated with a single statistical approach a single ranked list of differentially expressed proteins is assembled by comparing edgeR and DESeq q-values directly with the false discovery rate (FDR) calculated by baySeq. This statistical approach is then extended when applied to spectral count data derived from multiple proteomic pipelines. The individual statistical results from multiple proteomic pipelines are integrated and cross-validated by means of collapsing protein groups.