Missing value imputation in high-dimensional phenomic data: imputable or not, and how?

Missing value imputation in high-dimensional phenomic data: imputable or not, and how?
复制标题

DOI:
10.1186/s12859-014-0346-6
复制
发表时间:
2014-11-05
期刊:
影响因子:
3
通讯作者:
Tseng GC
Tseng GC
中科院分区:
生物学4区
文献类型:
--
作者:
Liao SG;Lin Y;Kang DD;Chandra D;Bon J;Kaminski N;Sciurba FC;Tseng GC

文献摘要

参考文献

被引文献

相似文献

在复杂疾病的现代生物医学研究中,经常收集大量的人口统计和临床变量,本文称为表组数据,并且在数据收集过程中不可避免地会出现缺失值(MV)。由于许多下游统计和生物信息学方法需要完整的数据矩阵,插补是一种常见且实用的解决方案。在微阵列实验等高通量实验中,测量连续强度,许多成熟的缺失值插补方法已经发展并广泛应用。已经开发了许多用于微阵列数据缺失数据插补的方法。然而,大型表组数据包含连续、名义、二进制和序数数据类型,这使得大多数方法无法应用。尽管过去几年已经开发了几种方法,但对于表组缺失数据插补,还没有提出一个完整的指南。在本文中,我们研究了现有的表型数据插补方法,提出了一种自训练选择(STS)方案来选择最佳插补方法,并为一般应用提供实用指南。我们引入了“可归责性度量”(IM)的新概念,以识别根本上不足以归责的缺失值。此外,我们还开发了 K 最近邻 (KNN) 方法的四种变体,并与链式方程多元插补 (MICE) 和 missForest 两种现有方法进行了比较。这四种变体是按变量插补 (KNN-V)、按受试者插补 (KNN-S)、加权混合插补 (KNN-H) 和自适应加权混合插补 (KNN-A)。我们对三个肺部疾病表组数据集进行了模拟并应用了不同的插补方法和 STS 方案来评估这些方法。 R 包“phenomeImpute”已公开发布。模拟和对真实数据集的应用表明,MICE 通常表现不佳;尽管没有一种方法普遍表现最好,但 KNN-A、KNN-H 和随机森林属于表现最好的方法。具有低可归因性的缺失值的归因会大大增加归因误差,并可能会恶化下游分析。 STS方案通过评估第二层缺失模拟中的方法来准确地选择最佳方法。所有模拟和真实数据分析的源文件都可以在作者的出版网站上找到。本文的在线版本 (doi:10.1186/s12859-014-0346-6) 包含补充材料,可供授权用户使用。
In modern biomedical research of complex diseases, a large number of demographic and clinical variables, herein called phenomic data, are often collected and missing values (MVs) are inevitable in the data collection process. Since many downstream statistical and bioinformatics methods require complete data matrix, imputation is a common and practical solution. In high-throughput experiments such as microarray experiments, continuous intensities are measured and many mature missing value imputation methods have been developed and widely applied. Numerous methods for missing data imputation of microarray data have been developed. Large phenomic data, however, contain continuous, nominal, binary and ordinal data types, which void application of most methods. Though several methods have been developed in the past few years, not a single complete guideline is proposed with respect to phenomic missing data imputation. In this paper, we investigated existing imputation methods for phenomic data, proposed a self-training selection (STS) scheme to select the best imputation method and provide a practical guideline for general applications. We introduced a novel concept of “imputability measure” (IM) to identify missing values that are fundamentally inadequate to impute. In addition, we also developed four variations of K-nearest-neighbor (KNN) methods and compared with two existing methods, multivariate imputation by chained equations (MICE) and missForest. The four variations are imputation by variables (KNN-V), by subjects (KNN-S), their weighted hybrid (KNN-H) and an adaptively weighted hybrid (KNN-A). We performed simulations and applied different imputation methods and the STS scheme to three lung disease phenomic datasets to evaluate the methods. An R package “phenomeImpute” is made publicly available. Simulations and applications to real datasets showed that MICE often did not perform well; KNN-A, KNN-H and random forest were among the top performers although no method universally performed the best. Imputation of missing values with low imputability measures increased imputation errors greatly and could potentially deteriorate downstream analyses. The STS scheme was accurate in selecting the optimal method by evaluating methods in a second layer of missingness simulation. All source files for the simulation and the real data analyses are available on the author’s publication website. The online version of this article (doi:10.1186/s12859-014-0346-6) contains supplementary material, which is available to authorized users.
DOI: 10.1002/sim.2939
发表时间: 2008-01-15
影响因子: 2
作者:
Little, Roderick J.;Yosef, Matheos;Harlow, Sioban D.
通讯作者: Harlow, Sioban D.
DOI: 10.1006/jmva.1995.1029
发表时间: 1995-04-01
影响因子: 1.6
作者:
LIU, C
通讯作者: LIU, C
DOI: 10.1126/science.29.751.823
发表时间: 1909-01-01
期刊: SCIENCE
影响因子: 56.9
作者:
Boas, F
通讯作者: Boas, F
DOI: 10.1093/bioinformatics/btq613
发表时间: 2011-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Oh, Sunghee;Kang, Dongwan D.;Tseng, George C.
通讯作者: Tseng, George C.
DOI: 10.1093/nar/gnh026
发表时间: 2004-02-01
影响因子: 14.9
作者:
Bo, TH;Dysvik, J;Jonassen, I
通讯作者: Jonassen, I