Identifying differentially expressed proteins in two-dimensional electrophoresis experiments: inputs from transcriptomics statistical tools

Identifying differentially expressed proteins in two-dimensional electrophoresis experiments: inputs from transcriptomics statistical tools
复制标题

DOI:
10.1093/bioinformatics/btt464
复制
发表时间:
2013-11-01
期刊:
影响因子:
5.8
通讯作者:
Pichereau, Vianney
Pichereau, Vianney
中科院分区:
生物学3区
文献类型:
--
作者:
Artigaud, Sebastien;Gauthier, Olivier;Pichereau, Vianney

文献摘要

被引文献

相似文献

背景:双向电泳是蛋白质组学中的一种重要方法,可以表征蛋白质的功能和表达。这通常意味着在两种对比条件之间差异表达的蛋白质的鉴定,例如,人类蛋白质组学生物标志物发现中的健康与患病,以及动物实验中的应激条件与对照。导致这种鉴定的统计程序是2-DE分析工作流程中的关键步骤。它们包括一个标准化步骤和一个测试和概率校正多个测试。由数据的高维性和大规模多重检验引起的统计问题在转录组学中比蛋白质组学更活跃,特别是在微阵列分析中。因此,我们建议适应创新的统计工具开发的微阵列分析,并将它们纳入2-DE分析pipeline.Results:在这篇文章中,我们评估了不同的归一化程序,不同的统计测试和错误发现率计算方法与真实的和模拟数据集的性能。我们证明,使用的统计程序改编自微阵列导致显着增加的权力,以及最小化的假阳性发现率。更具体地说,我们获得了最好的结果,在可靠性和灵敏度方面,使用'温和的t检验'从史密斯与经典的错误发现率从Benjamini和Hochberg。
Background: Two-dimensional electrophoresis is a crucial method in proteomics that allows the characterization of proteins' function and expression. This usually implies the identification of proteins that are differentially expressed between two contrasting conditions, for example, healthy versus diseased in human proteomics biomarker discovery and stressful conditions versus control in animal experimentation. The statistical procedures that lead to such identifications are critical steps in the 2-DE analysis workflow. They include a normalization step and a test and probability correction for multiple testing. Statistical issues caused by the high dimensionality of the data and large-scale multiple testing have been a more active topic in transcriptomics than proteomics, especially in microarray analysis. We thus propose to adapt innovative statistical tools developed for microarray analysis and incorporate them in the 2-DE analysis pipeline.Results: In this article, we evaluate the performance of different normalization procedures, different statistical tests and false discovery rate calculation methods with both real and simulated datasets. We demonstrate that the use of statistical procedures adapted from microarrays lead to notable increase in power as well as a minimization of false-positive discovery rate. More specifically, we obtained the best results in terms of reliability and sensibility when using the 'moderate t-test' from Smyth in association with classic false discovery rate from Benjamini and Hochberg.