Improving cell mixture deconvolution by identifying optimal DNA methylation libraries (IDOL).

Improving cell mixture deconvolution by identifying optimal DNA methylation libraries (IDOL).
复制标题

DOI:
10.1186/s12859-016-0943-7
复制
发表时间:
2016-03-08
期刊:
影响因子:
3
通讯作者:
Kelsey KT
Kelsey KT
中科院分区:
生物学4区
文献类型:
--
作者:
Koestler DC;Jones MJ;Usset J;Christensen BC;Butler RA;Kobor MS;Wiencke JK;Kelsey KT

文献摘要

被引文献

相似文献

由于细胞异质性造成的混淆是目前表观全基因组关联研究(EWAS)面临的最大挑战之一。利用DNA甲基化的组织特异性对异质生物样本的细胞混合物进行去卷积的统计方法提供了一种有前景的解决方案,然而,这种方法的性能完全取决于用于去卷积的甲基化标记物库。在这里,我们介绍了一种新的算法,用于识别最佳库(IDOL),动态扫描一组候选的细胞特异性甲基化标记,以找到库,优化从细胞混合物反卷积获得的细胞分数估计的准确性。将IDOL应用于由具有全血DNA甲基化数据(Illumina HumanMethylation 450 BeadArray(HM 450))和细胞组成的流式细胞术测量的样品组成的训练集,揭示了由300个CpG位点组成的优化文库。当比较现有文库时,通过IDOL鉴定的文库显示出对整个免疫细胞景观的显著更好的总体区分(p = 0.038),并且导致15对白细胞亚型中的14对的区分改善。使用IDOL文库对训练集中的样品的细胞组成的估计与其各自的流式细胞术测量高度相关,所有细胞特异性R2>0.99,并且白细胞亚型的均方根误差(RMSE)范围为[0.97%至1.33%]。使用两个额外的HM 450数据集对优化的IDOL文库的独立验证显示出类似的强预测性能,所有细胞特异性R2>0.90且RMSE<4.00%。在模拟研究中,与竞争文库相比,使用IDOL文库对细胞组成进行的调整导致了一致较低的假阳性率,同时也证明了在两个大型公开可用的HM 450数据集内解释DNA甲基化的表观基因组范围内变异的能力有所提高。尽管与用于全血混合物去卷积的现有文库相比,由一半的CpG组成,但本文鉴定的优化的IDOL文库在所有考虑的数据集上产生了出色的预测性能,并证明了改善涉及细胞分布调整的EWAS的操作特征的潜力。除了为EWAS社区提供用于全血混合物去卷积的优化库之外,我们的工作还为提高细胞混合物去卷积的准确性的库的组装建立了系统的和可推广的框架。本文的在线版本(doi:10.1186/s12859-016-0943-7)包含补充材料,可供授权用户使用。
Confounding due to cellular heterogeneity represents one of the foremost challenges currently facing Epigenome-Wide Association Studies (EWAS). Statistical methods leveraging the tissue-specificity of DNA methylation for deconvoluting the cellular mixture of heterogenous biospecimens offer a promising solution, however the performance of such methods depends entirely on the library of methylation markers being used for deconvolution. Here, we introduce a novel algorithm for Identifying Optimal Libraries (IDOL) that dynamically scans a candidate set of cell-specific methylation markers to find libraries that optimize the accuracy of cell fraction estimates obtained from cell mixture deconvolution. Application of IDOL to training set consisting of samples with both whole-blood DNA methylation data (Illumina HumanMethylation450 BeadArray (HM450)) and flow cytometry measurements of cell composition revealed an optimized library comprised of 300 CpG sites. When compared existing libraries, the library identified by IDOL demonstrated significantly better overall discrimination of the entire immune cell landscape (p = 0.038), and resulted in improved discrimination of 14 out of the 15 pairs of leukocyte subtypes. Estimates of cell composition across the samples in the training set using the IDOL library were highly correlated with their respective flow cytometry measurements, with all cell-specific R2>0.99 and root mean square errors (RMSEs) ranging from [0.97 % to 1.33 %] across leukocyte subtypes. Independent validation of the optimized IDOL library using two additional HM450 data sets showed similarly strong prediction performance, with all cell-specific R2>0.90 and RMSE<4.00 %. In simulation studies, adjustments for cell composition using the IDOL library resulted in uniformly lower false positive rates compared to competing libraries, while also demonstrating an improved capacity to explain epigenome-wide variation in DNA methylation within two large publicly available HM450 data sets. Despite consisting of half as many CpGs compared to existing libraries for whole blood mixture deconvolution, the optimized IDOL library identified herein resulted in outstanding prediction performance across all considered data sets and demonstrated potential to improve the operating characteristics of EWAS involving adjustments for cell distribution. In addition to providing the EWAS community with an optimized library for whole blood mixture deconvolution, our work establishes a systematic and generalizable framework for the assembly of libraries that improve the accuracy of cell mixture deconvolution. The online version of this article (doi:10.1186/s12859-016-0943-7) contains supplementary material, which is available to authorized users.