LSimpute: accurate estimation of missing values in microarray data with least squares methods

LSimpute: accurate estimation of missing values in microarray data with least squares methods
复制标题

DOI:
10.1093/nar/gnh026
复制
发表时间:
2004-02-01
影响因子:
14.9
通讯作者:
Jonassen, I
Jonassen, I
中科院分区:
生物学2区
文献类型:
--
作者:
Bo, TH;Dysvik, J;Jonassen, I

文献摘要

被引文献

相似文献

微阵列实验产生数据集,其中包含一组生物样本中数千个基因表达水平的信息。不幸的是,这样的实验往往会产生多个缺失的表达值,通常是由于各种实验问题。由于许多基因表达分析算法需要一个完整的数据矩阵作为输入,为了分析可用的数据,必须估计缺失的值。或者,可以移除基因和阵列,直到没有缺失值。然而,对于只有少量缺失值的基因或阵列,需要对这些值进行估算。为了使后续分析尽可能提供信息,对缺失基因表达值的估计是准确的是至关重要的。对于聚类方法(如分层聚类或K-means聚类)来说,数据中少量估计错误的缺失值可能足以产生误导性的结果。因此,需要精确的缺失值估计方法。我们提出了基于最小二乘原理的微阵列数据集中缺失值估计的新方法,并利用基因和阵列之间的相关性。对于这组方法,我们使用通用的引用名称LSimpute。我们通过随机剔除数据(标记为缺失),将我们的方法与广泛使用的KNNimpute在公共数据集的三个完整数据矩阵上的估计精度进行了比较。从这些测试中,我们得出结论,我们的LSimpute方法产生的估计值始终比使用KNNimpute获得的估计值更准确。此外,我们研究了一种基于期望最大化(EM)的更经典的缺失值估计方法。我们将我们的电磁实现称为EMimpute,并将使用EMimpute方法的估计误差与我们的新方法产生的估计误差进行了比较。结果表明,平均而言,我们最好的LSimpute方法的估计至少与最好的EMimpute算法的估计一样准确。
Microarray experiments generate data sets with information on the expression levels of thousands of genes in a set of biological samples. Unfortunately, such experiments often produce multiple missing expression values, normally due to various experimental problems. As many algorithms for gene expression analysis require a complete data matrix as input, the missing values have to be estimated in order to analyze the available data. Alternatively, genes and arrays can be removed until no missing values remain. However, for genes or arrays with only a small number of missing values, it is desirable to impute those values. For the subsequent analysis to be as informative as possible, it is essential that the estimates for the missing gene expression values are accurate. A small amount of badly estimated missing values in the data might be enough for clustering methods, such as hierachical clustering or K-means clustering, to produce misleading results. Thus, accurate methods for missing value estimation are needed. We present novel methods for estimation of missing values in microarray data sets that are based on the least squares principle, and that utilize correlations between both genes and arrays. For this set of methods, we use the common reference name LSimpute. We compare the estimation accuracy of our methods with the widely used KNNimpute on three complete data matrices from public data sets by randomly knocking out data (labeling as missing). From these tests, we conclude that our LSimpute methods produce estimates that consistently are more accurate than those obtained using KNNimpute. Additionally, we examine a more classic approach to missing value estimation based on expectation maximization (EM). We refer to our EM implementations as EMimpute, and the estimate errors using the EMimpute methods are compared with those our novel methods produce. The results indicate that on average, the estimates from our best performing LSimpute method are at least as accurate as those from the best EMimpute algorithm.