A comparison of feature selection and classification methods in DNA methylation studies using the Illumina Infinium platform.

A comparison of feature selection and classification methods in DNA methylation studies using the Illumina Infinium platform.
复制标题

DOI:
10.1186/1471-2105-13-59
复制
发表时间:
2012-04-24
期刊:
影响因子:
3
通讯作者:
Teschendorff AE
Teschendorff AE
中科院分区:
生物学4区
文献类型:
--
作者:
Zhuang J;Widschwendter M;Teschendorff AE

文献摘要

被引文献

相似文献

27K Illumina Infinium甲基化微珠芯片是一项流行的高通量技术,可以检测超过27,000个CPG的甲基化状态。虽然已经在基因表达数据的背景下全面探索了特征选择和分类方法,但相对较少知道如何在Illumina Infinium甲基化数据的背景下最好地执行特征选择或分类。鉴于表观基因组学在癌症和其他复杂遗传疾病中的重要性日益上升,以及即将进行的表观基因组广泛关联研究,确定在这一新背景下提供改进推断的统计方法至关重要。使用总共7个大型Illumina Infinium27k甲基化数据集,包括来自广泛组织的1,000多个样本,我们在这里提供了对DNA甲基化数据的流行特征选择、降维和分类方法的评估。具体地说,我们评估了方差滤波、有监督主成分(SPCA)和DNA甲基化量化指标的选择对下游统计推断的影响。我们表明,对于相对较大的样本量,使用检验统计量的特征选择对于M值和β值是相似的,但在小样本量的限制下,M值允许更可靠地识别真正的阳性。我们还表明,方差滤波对特征选择的影响是研究特定的,并且依赖于感兴趣的表型和所描述的组织类型。具体地说,我们发现,在效应大的研究中,方差滤波提高了对真阳性的检测,但在效应小但显著的研究中,它可能会导致较差的性能。相比之下,有监督的主成分提高了统计能力,特别是在效应规模较小的研究中。我们还证明了使用弹性网络和支持向量机(SVM)的分类方法明显优于LASSO和SPCA等竞争方法。最后,在癌症诊断的非监督模型中,我们发现非负矩阵分解(NMF)明显优于主成分分析。我们的结果强调了根据DNA甲基化研究的样本大小和生物学背景调整特征选择和分类方法的重要性。弹性网络作为一种强大的分类算法出现在大规模DNA甲基化研究中,而NMF在非监督环境下表现良好。这里提出的见解将对任何开始使用Illumina Infinium珠阵进行大规模DNA甲基化分析的研究有用。
The 27k Illumina Infinium Methylation Beadchip is a popular high-throughput technology that allows the methylation state of over 27,000 CpGs to be assayed. While feature selection and classification methods have been comprehensively explored in the context of gene expression data, relatively little is known as to how best to perform feature selection or classification in the context of Illumina Infinium methylation data. Given the rising importance of epigenomics in cancer and other complex genetic diseases, and in view of the upcoming epigenome wide association studies, it is critical to identify the statistical methods that offer improved inference in this novel context. Using a total of 7 large Illumina Infinium 27k Methylation data sets, encompassing over 1,000 samples from a wide range of tissues, we here provide an evaluation of popular feature selection, dimensional reduction and classification methods on DNA methylation data. Specifically, we evaluate the effects of variance filtering, supervised principal components (SPCA) and the choice of DNA methylation quantification measure on downstream statistical inference. We show that for relatively large sample sizes feature selection using test statistics is similar for M and β-values, but that in the limit of small sample sizes, M-values allow more reliable identification of true positives. We also show that the effect of variance filtering on feature selection is study-specific and dependent on the phenotype of interest and tissue type profiled. Specifically, we find that variance filtering improves the detection of true positives in studies with large effect sizes, but that it may lead to worse performance in studies with smaller yet significant effect sizes. In contrast, supervised principal components improves the statistical power, especially in studies with small effect sizes. We also demonstrate that classification using the Elastic Net and Support Vector Machine (SVM) clearly outperforms competing methods like LASSO and SPCA. Finally, in unsupervised modelling of cancer diagnosis, we find that non-negative matrix factorisation (NMF) clearly outperforms principal components analysis. Our results highlight the importance of tailoring the feature selection and classification methodology to the sample size and biological context of the DNA methylation study. The Elastic Net emerges as a powerful classification algorithm for large-scale DNA methylation studies, while NMF does well in the unsupervised context. The insights presented here will be useful to any study embarking on large-scale DNA methylation profiling using Illumina Infinium beadarrays.