Missing value estimation methods for DNA microarrays

Missing value estimation methods for DNA microarrays
复制标题

DOI:
10.1093/bioinformatics/17.6.520
复制
发表时间:
2001-06-01
期刊:
影响因子:
5.8
通讯作者:
Altman, RB
Altman, RB
中科院分区:
生物学3区
文献类型:
--
作者:
Troyanskaya, O;Cantor, M;Altman, RB

文献摘要

被引文献

相似文献

动机:基因表达微阵列实验可以生成具有多个缺少表达值的数据集。不幸的是,许多用于基因表达分析的算法都需要一个基因阵列值的完整矩阵作为输入。例如,诸如层次聚类和K-均值聚类之类的方法对于丢失数据并不强大,即使有一些丢失的值,也可能会失去有效性。因此,需要用于最小化数据集对分析的影响的影响,并增加可以应用这些算法的数据集范围,以最大程度地限制数据集对分析的影响。在本报告中,我们研究了用于估计数据丢失的自动化方法。回报:我们对基因微阵列数据中缺失值的几种方法进行了比较研究。我们实施并评估了三种方法:基于奇异的值分解(SVD)方法(SVDIMPUTE),加权K-Nearest邻居(Knnimpute)和行平均值。我们使用各种参数设置和不同的实际数据集评估了方法,并评估了归纳方法的鲁棒性,以在1-20%缺失值范围内丢失数据量。我们表明,Knnimpute似乎比SVDImpute提供了一种更强大和更敏感的方法,用于缺少价值估计,并且SVDIMPUTE和KNNIMPTUTE都超过了常用的行平均方法LAS,可以用零填充缺失值)。我们报告了比较实验的结果,并提供了建议和工具,以准确估计各种条件下缺失的微阵列数据。
Motivation: Gene expression microarray experiments can generate data sets with multiple missing expression values. Unfortunately, many algorithms for gene expression analysis require a complete matrix of gene array values as input. For example, methods such as hierarchical clustering and K-means clustering are not robust to missing data, and may lose effectiveness even with a few missing values. Methods for imputing missing data are needed, therefore, to minimize the effect of incomplete data sets on analyses, and to increase the range of data sets to which these algorithms can be applied. In this report, we investigate automated methods for estimating missing data.Results: We present a comparative study of several methods for the estimation of missing values in gene microarray data. We implemented and evaluated three methods: a Singular Value Decomposition (SVD) based method (SVDimpute), weighted K-nearest neighbors (KNNimpute), and row average. We evaluated the methods using a variety of parameter settings and over different real data sets, and assessed the robustness of the imputation methods to the amount of missing data over the range of 1-20% missing values. We show that KNNimpute appears to provide a more robust and sensitive method for missing value estimation than SVDimpute, and both SVDimpute and KNNimpute surpass the commonly used row average method las well as filling missing values with zeros). We report results of the comparative experiments and provide recommendations and tools for accurate estimation of missing microarray data under a variety of conditions.