Using multivariate mixed-effects selection models for analyzing batch-processed proteomics data with non-ignorable missingness.

Using multivariate mixed-effects selection models for analyzing batch-processed proteomics data with non-ignorable missingness.
复制标题

DOI:
10.1093/biostatistics/kxy022
复制
发表时间:
2018-06
期刊:
影响因子:
2.1
通讯作者:
Jiebiao Wang;Pei Wang;D. Hedeker;Lin S. Chen
Jiebiao Wang;Pei Wang;D. Hedeker;Lin S. Chen
中科院分区:
数学2区
文献类型:
--
作者:
Jiebiao Wang;Pei Wang;D. Hedeker;Lin S. Chen

文献摘要

相似文献

在定量蛋白质组学中,质量标签标记技术已被广泛应用于质谱学实验。这些技术允许在一次实验中检测和定量一批多个样品中的多肽(短氨基酸序列)和蛋白质,从而极大地提高了蛋白质图谱的效率。然而,样品的批处理也会导致严重的批次效应和在批次级别发生的不可忽视的数据丢失。受临床蛋白质组肿瘤分析联盟提供的乳腺癌蛋白质组数据的启发,考虑批次效应和不可忽略的缺失性,我们开发了两个定制的多变量混合效应选择模型(MvMISE),用于联合分析标记蛋白质组数据中的多个相关多肽/蛋白质。通过采用多变量方法,我们可以借用同一蛋白质的多个多肽或同一生物途径中的多个蛋白质的信息,从而获得更好的统计效率和生物学解释。这两种不同的模型解释了一组多肽或蛋白质之间不同的关联结构。具体地说,为了从相同的蛋白质中模拟多个多肽,我们使用了因子分析随机效应结构来表征多肽之间的高度和相似的相关性。为了模拟功能通路中多个蛋白质之间的生物依赖关系,我们引入了误差精度矩阵的图解套索惩罚,并实现了一种基于乘子交替方向法的高效算法。仿真结果验证了所提模型的优越性。将所提出的方法应用于激励数据集,我们识别出在三阴性乳腺肿瘤中表现出与其他乳腺肿瘤不同的活动模式的磷蛋白和生物通路。所提出的方法也可应用于其他基于具有或不具有不可忽略缺失的聚类数据的高维多变量分析。
In quantitative proteomics, mass tag labeling techniques have been widely adopted in mass spectrometry experiments. These techniques allow peptides (short amino acid sequences) and proteins from multiple samples of a batch being detected and quantified in a single experiment, and as such greatly improve the efficiency of protein profiling. However, the batch-processing of samples also results in severe batch effects and non-ignorable missing data occurring at the batch level. Motivated by the breast cancer proteomic data from the Clinical Proteomic Tumor Analysis Consortium, in this work, we developed two tailored multivariate MIxed-effects SElection models (mvMISE) to jointly analyze multiple correlated peptides/proteins in labeled proteomics data, considering the batch effects and the non-ignorable missingness. By taking a multivariate approach, we can borrow information across multiple peptides of the same protein or multiple proteins from the same biological pathway, and thus achieve better statistical efficiency and biological interpretation. These two different models account for different correlation structures among a group of peptides or proteins. Specifically, to model multiple peptides from the same protein, we employed a factor-analytic random effects structure to characterize the high and similar correlations among peptides. To model biological dependence among multiple proteins in a functional pathway, we introduced a graphical lasso penalty on the error precision matrix, and implemented an efficient algorithm based on the alternating direction method of multipliers. Simulations demonstrated the advantages of the proposed models. Applying the proposed methods to the motivating data set, we identified phosphoproteins and biological pathways that showed different activity patterns in triple negative breast tumors versus other breast tumors. The proposed methods can also be applied to other high-dimensional multivariate analyses based on clustered data with or without non-ignorable missingness.