Improving cluster-based missing value estimation of DNA microarray data

Improving cluster-based missing value estimation of DNA microarray data
复制标题

DOI:
10.1016/j.bioeng.2007.04.003
复制
发表时间:
2007-06-01
期刊:
BIOMOLECULAR ENGINEERING
影响因子:
--
通讯作者:
Menezes, Jose C.
Menezes, Jose C.
中科院分区:
其他
文献类型:
--
作者:
Bras, Ligia P.;Menezes, Jose C.

文献摘要

被引文献

相似文献

我们提出了一种改进的加权K-最近邻填补方法(KNNimpute)的缺失值(MV)估计在微阵列数据的基础上估计数据的重用。该方法称为迭代KNN插补(IKNNimpute),因为使用最近的估计值迭代地进行估计,在不同条件下评估了IKNNimpme的估计效率(数据类型,缺失数据的分数和结构)通过归一化均方根误差(NRMSE)和估计值与真实值之间的相关系数,并与其他基于聚类的估计方法(KNNimpute和序贯KNN)进行了比较。我们通过检查MV估计后丢失的差异表达基因,进一步研究了插补对使用SAM检测差异表达基因的影响。性能指标给出了一致的结果,表明lKNNimpute的迭代过程可以增强基于聚类的方法在高缺失率下的预测能力,在非时间序列实验中以及在包括时间序列和非时间序列数据的数据集中,因为具有MV的基因的信息被更有效地使用,并且迭代过程允许细化MV估计。更重要的是,IKNN对差异表达基因的检测具有较小的不利影响。(c)2007 Elsevier B. V.保留所有权利。
We present a modification of the weighted K-nearest neighbours imputation method (KNNimpute) for missing values (MVs) estimation in microarray data based on the reuse of estimated data. The method was called iterative KNN imputation (IKNNimpute) as the estimation is performed iteratively using the recently estimated values.The estimation efficiency of lKNNimpme was assessed under different conditions (data type, fraction and structure of missing data) by the normalized root mean squared error (NRMSE) and the correlation coefficients between estimated and true values, and compared with that of other cluster-based estimation methods (KNNimpute and sequential KNN). We further investigated the influence of imputation on the detection of differentially expressed genes using SAM by examining the differentially expressed genes that are lost after MV estimation.The performance measures give consistent results, indicating that the iterative procedure of lKNNimpute can enhance the prediction ability of cluster-based methods in the presence of high missing rates, in non-time series experiments and in data sets comprising both time series and non-time series data, because the information of the genes having MVs is used more efficiently and the iterative procedure allows refining the MV estimates. More importantly, IKNN has a smaller detrimental effect on the detection of differentially expressed genes. (c) 2007 Elsevier B.V. All rights reserved.