Guidance for DNA methylation studies: statistical insights from the Illumina EPIC array

Guidance for DNA methylation studies: statistical insights from the Illumina EPIC array
复制标题

DOI:
10.1186/s12864-019-5761-7
复制
发表时间:
2019-05-14
期刊:
影响因子:
4.4
通讯作者:
Hannon, Eilis
Hannon, Eilis
中科院分区:
生物学2区
文献类型:
--
作者:
Mansell, Georgina;Gorrie-Stone, Tyler J.;Hannon, Eilis

文献摘要

被引文献

相似文献

背景:旨在鉴定与复杂表型相关的DNA甲基化差异的研究数量稳步增加。表观遗传流行病学在研究设计和解释方面的许多挑战已经得到了详细的讨论,但是有一些分析方面的问题是突出的,需要进一步探索。在这项研究中,我们试图解决三个分析问题。首先,我们量化了多重测试负担,并提出了一个标准的统计显著性阈值,用于识别与结果相关的DNA甲基化位点。其次,我们确定线性回归(大多数研究选择的统计工具)是否合适,以及它是否受到DNA甲基化数据的潜在分布的影响。最后,我们评估了足够动力的DNA甲基化关联研究所需的样本量。结果:我们使用Illumina EPIC阵列评估DNA甲基化关联分析的统计特性,在一项基于大规模人群的研究——Understanding Society队列(n=1175)中对DNA甲基化进行了量化。通过模拟零DNA甲基化研究,我们生成了随机预期的p值分布,并计算出EPIC阵列研究的5%家庭误差为9x10(-8)。接下来,我们检验了DNA甲基化数据是否违反了线性回归的假设,发现大多数位点不满足正态残差的假设。然而,我们没有发现证据表明这种偏倚通过增加受影响部位为假阳性的可能性来影响分析。最后,我们对基于EPIC的DNA甲基化研究进行了功率计算,结果表明,现有的研究数据接近1000个样本,足以检测大多数位点的微小差异。结论提出了P
BackgroundThere has been a steady increase in the number of studies aiming to identify DNA methylation differences associated with complex phenotypes. Many of the challenges of epigenetic epidemiology regarding study design and interpretation have been discussed in detail, however there are analytical concerns that are outstanding and require further exploration. In this study we seek to address three analytical issues. First, we quantify the multiple testing burden and propose a standard statistical significance threshold for identifying DNA methylation sites that are associated with an outcome. Second, we establish whether linear regression, the chosen statistical tool for the majority of studies, is appropriate and whether it is biased by the underlying distribution of DNA methylation data. Finally, we assess the sample size required for adequately powered DNA methylation association studies.ResultsWe quantified DNA methylation in the Understanding Society cohort (n=1175), a large population based study, using the Illumina EPIC array to assess the statistical properties of DNA methylation association analyses. By simulating null DNA methylation studies, we generated the distribution of p-values expected by chance and calculated the 5% family-wise error for EPIC array studies to be 9x10(-8). Next, we tested whether the assumptions of linear regression are violated by DNA methylation data and found that the majority of sites do not satisfy the assumption of normal residuals. Nevertheless, we found no evidence that this bias influences analyses by increasing the likelihood of affected sites to be false positives. Finally, we performed power calculations for EPIC based DNA methylation studies, demonstrating that existing studies with data on similar to 1000 samples are adequately powered to detect small differences at the majority of sites.ConclusionWe propose that a significance threshold of P