Improved quality control processing of peptide-centric LC-MS proteomics data.

Improved quality control processing of peptide-centric LC-MS proteomics data.
复制标题

DOI:
10.1093/bioinformatics/btr479
复制
发表时间:
2011-10-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Webb-Robertson BJ
Webb-Robertson BJ
中科院分区:
其他
文献类型:
--
作者:
Matzke MM;Waters KM;Metz TO;Jacobs JM;Sims AC;Baric RS;Pounds JG;Webb-Robertson BJ

文献摘要

参考文献

被引文献

相似文献

动机:在差异肽峰强度(即丰度测量)的分析中,具有低质量肽丰度数据的LC-MS分析可能会使下游统计分析产生偏差,因此会使高质量数据集的生物学解释产生偏差。尽管在光谱处理方面已经做出了相当大的努力来确保肽鉴定的质量,但迄今为止,后续肽丰度数据矩阵的质量评估仅限于逐批相关性或单个肽组分的主观目视检查。识别统计异常值是蛋白质组学数据处理中的关键步骤,因为许多下游统计分析[例如方差分析(ANOVA)]依赖于对样本方差的准确估计,并且其结果受到极值的影响。结果:我们描述了一种新的多元统计策略,用于识别具有极端肽丰度分布的LC-MS运行。与当前方法(运行相关性)的比较表明,通过多变量策略识别离群运行的速度明显更好。模拟研究还表明,这种策略显着优于相关性单独在统计上极端的液相色谱-质谱(LC-MS)运行的识别。可用性:https://www.biopilot.org/docs/Software/RMD.php联系:bj@pnl.gov补充信息:补充材料可在生物信息学在线。
Motivation: In the analysis of differential peptide peak intensities (i.e. abundance measures), LC-MS analyses with poor quality peptide abundance data can bias downstream statistical analyses and hence the biological interpretation for an otherwise high-quality dataset. Although considerable effort has been placed on assuring the quality of the peptide identification with respect to spectral processing, to date quality assessment of the subsequent peptide abundance data matrix has been limited to a subjective visual inspection of run-by-run correlation or individual peptide components. Identifying statistical outliers is a critical step in the processing of proteomics data as many of the downstream statistical analyses [e.g. analysis of variance (ANOVA)] rely upon accurate estimates of sample variance, and their results are influenced by extreme values. Results: We describe a novel multivariate statistical strategy for the identification of LC-MS runs with extreme peptide abundance distributions. Comparison with current method (run-by-run correlation) demonstrates a significantly better rate of identification of outlier runs by the multivariate strategy. Simulation studies also suggest that this strategy significantly outperforms correlation alone in the identification of statistically extreme liquid chromatography-mass spectrometry (LC-MS) runs. Availability: https://www.biopilot.org/docs/Software/RMD.php Contact: bj@pnl.gov Supplementary information: Supplementary material is available at Bioinformatics online.
DOI: 10.1186/1756-0381-2-4
发表时间: 2009-04-07
期刊: BioData mining
影响因子: 4.5
作者:
Schulz-Trieglaff O;Machtejevas E;Reinert K;Schlüter H;Thiemann J;Unger K
通讯作者: Unger K
DOI: 10.1093/bioinformatics/btm281
发表时间: 2007-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Monroe, Matthew E.;Tolic, Nikola;Smith, Richard D.
通讯作者: Smith, Richard D.
DOI: 10.1021/ac034790h
发表时间: 2003-12-15
影响因子: 7.4
作者:
MacCoss, MJ;Wu, CC;Yates, JR
通讯作者: Yates, JR
DOI: 10.2307/2291724
发表时间: 1996-09-01
影响因子: 3.7
作者:
Rocke, DM;Woodruff, DL
通讯作者: Woodruff, DL
DOI: 10.1016/j.clinbiochem.2010.04.071
发表时间: 2010-08-01
影响因子: 2.8
作者:
Jain, Ram B.
通讯作者: Jain, Ram B.