MatchMixeR: a cross-platform normalization method for gene expression data integration.

MatchMixeR: a cross-platform normalization method for gene expression data integration.
复制标题

MatchMixeR:一种用于基因表达数据集成的跨平台标准化方法。

DOI:
10.1093/bioinformatics/btz974
复制
发表时间:
2020
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Zhang,Jinfeng
Zhang,Jinfeng
中科院分区:
--
文献类型:
--
作者:
Zhang,Serin;Shao,Jiang;Yu,Disa;Qiu,Xing;Zhang,Jinfeng

文献摘要

相似文献

运动结合从不同平台产生的基因表达(GE)谱使以前由于样本量限制而无法进行的研究成为可能。已经开发了几种跨平台归一化方法来消除平台之间的系统差异,但它们也可以消除数据集之间有意义的生物学差异。在这项工作中,我们提出了一种新的方法,消除了平台,而不是生物差异。我们被称为‘MatchMixeR’,我们通过线性混合效应回归(LMER)模型对平台差异进行建模,并根据在不同平台上测量的相同细胞系或组织的匹配GE谱来估计它们。然后,可以使用生成的模型来消除其他数据集中的平台差异。通过使用LMER,我们在参数估计中实现了更好的偏差-方差折衷。我们还设计了一种基于矩方法的计算效率较高的算法,该算法非常适合于超高维LMER分析。结果与几种重要的竞争方法相比,MatchMixeR获得了最高的归一化一致性。基于从不同平台集成的数据集的后续差异表达分析表明,使用MatchMixeR实现了真假发现之间的最佳权衡,并且这一优势在样本有限或组比例不平衡的数据集中更加明显。可用性和实现我们的方法在R包‘MatchMixeR’中实现,该R包可在https://github.com/dy16b/Cross-Platform-Normalization.Supplementary上免费获得信息补充数据可在BioInformation Online上获得。
MotivationCombining gene expression (GE) profiles generated from different platforms enables previously infeasible studies due to sample size limitations. Several cross-platform normalization methods have been developed to remove the systematic differences between platforms, but they may also remove meaningful biological differences among datasets. In this work, we propose a novel approach that removes the platform, not the biological differences. Dubbed as ‘MatchMixeR’, we model platform differences by a linear mixed effects regression (LMER) model, and estimate them from matched GE profiles of the same cell line or tissue measured on different platforms. The resulting model can then be used to remove platform differences in other datasets. By using LMER, we achieve better bias-variance trade-off in parameter estimation. We also design a computationally efficient algorithm based on the moment method, which is ideal for ultra-high-dimensional LMER analysis.ResultsCompared with several prominent competing methods, MatchMixeR achieved the highest after-normalization concordance. Subsequent differential expression analyses based on datasets integrated from different platforms showed that using MatchMixeR achieved the best trade-off between true and false discoveries, and this advantage is more apparent in datasets with limited samples or unbalanced group proportions.Availability and implementationOur method is implemented in a R-package, ‘MatchMixeR’, freely available at: https://github.com/dy16b/Cross-Platform-Normalization.Supplementary informationSupplementary data are available atBioinformaticsonline.