MSstatsTMT: Statistical Detection of Differentially Abundant Proteins in Experiments with Isobaric Labeling and Multiple Mixtures.

MSstatsTMT: Statistical Detection of Differentially Abundant Proteins in Experiments with Isobaric Labeling and Multiple Mixtures.
复制标题

DOI:
10.1074/mcp.ra120.002105
复制
发表时间:
2020-10
期刊:
Molecular & cellular proteomics : MCP
影响因子:
--
通讯作者:
Vitek O
Vitek O
中科院分区:
其他
文献类型:
--
作者:
Huang T;Choi M;Tzouros M;Golling S;Pandya NJ;Banfai B;Dunkley T;Vitek O

文献摘要

被引文献

相似文献

MSstatsTMT实现了相对蛋白质定量的一般统计方法,并在TMT标记的基于质谱的实验中测试差异丰度。它适用于多条件、多生物重复运行和多技术重复运行以及非平衡设计的实验。对受控混合物、模拟数据集和具有不同设计的三种生物调查的评估表明,在具有多种生物混合物的大规模实验中,MSstatsTMT平衡了检测差异丰度蛋白质的灵敏度和特异性。亮点TMT标记蛋白质组实验差异丰度分析的统计方法。适用于复杂或不平衡设计的大规模实验。一个开源的R/Bioconductor包,与流行的数据处理工具兼容。串联质量标签(TMT)是一种广泛应用于蛋白质组学研究的多重技术。它能够在单次MS运行中以高效率和高通量对多个生物样品中的蛋白质进行相对定量。然而,实验通常需要比单次运行所能容纳的更多的生物重复或条件,并且涉及多种TMT混合物和多次运行。这种大规模的实验结合了联合收割机的生物和技术变化的模式,是复杂的,独特的基于TMT的工作流程,并具有挑战性的下游统计分析。这些模式无法通过为其他技术(如无标记蛋白质组学或转录组学)设计的统计方法充分表征。本文提出了一种基于TMT标记的MS实验中蛋白质相对定量的一般统计方法。它适用于多条件、多生物重复运行和多技术重复运行以及非平衡设计的实验。它基于一系列灵活的线性混合效应模型,可以处理技术工件和缺失值的复杂模式。该方法在MSstatsTMT中实现,这是一个免费的开源R/Bioconductor包,与Proteome Discoverer,MaxQuant,OpenMS和SpectroMine等数据处理工具兼容。对受控混合物、模拟数据集和具有不同设计的三种生物调查的评估表明,在具有多种生物混合物的大规模实验中,MSstatsTMT平衡了检测差异丰度蛋白质的灵敏度和特异性。
MSstatsTMT implements a general statistical approach for relative protein quantification and tests for differential abundance in mass spectrometry-based experiments with TMT labeling. It is applicable to experiments with multiple conditions, multiple biological replicate runs and multiple technical replicate runs, and unbalanced designs. Evaluation on a controlled mixture, simulated datasets, and three biological investigations with diverse designs demonstrated that MSstatsTMT balanced the sensitivity and the specificity of detecting differentially abundant proteins, in large-scale experiments with multiple biological mixtures. Highlights Statistical approach for differential abundance analysis for proteomic experiments with TMT labeling. Applicable to large-scale experiments with complex or unbalanced design. An open-source R/Bioconductor package compatible with popular data processing tools. Tandem mass tag (TMT) is a multiplexing technology widely-used in proteomic research. It enables relative quantification of proteins from multiple biological samples in a single MS run with high efficiency and high throughput. However, experiments often require more biological replicates or conditions than can be accommodated by a single run, and involve multiple TMT mixtures and multiple runs. Such larger-scale experiments combine sources of biological and technical variation in patterns that are complex, unique to TMT-based workflows, and challenging for the downstream statistical analysis. These patterns cannot be adequately characterized by statistical methods designed for other technologies, such as label-free proteomics or transcriptomics. This manuscript proposes a general statistical approach for relative protein quantification in MS- based experiments with TMT labeling. It is applicable to experiments with multiple conditions, multiple biological replicate runs and multiple technical replicate runs, and unbalanced designs. It is based on a flexible family of linear mixed-effects models that handle complex patterns of technical artifacts and missing values. The approach is implemented in MSstatsTMT, a freely available open-source R/Bioconductor package compatible with data processing tools such as Proteome Discoverer, MaxQuant, OpenMS, and SpectroMine. Evaluation on a controlled mixture, simulated datasets, and three biological investigations with diverse designs demonstrated that MSstatsTMT balanced the sensitivity and the specificity of detecting differentially abundant proteins, in large-scale experiments with multiple biological mixtures.