Biological impact of missing-value imputation on downstream analyses of gene expression profiles

Biological impact of missing-value imputation on downstream analyses of gene expression profiles
复制标题

DOI:
10.1093/bioinformatics/btq613
复制
发表时间:
2011-01-01
期刊:
影响因子:
5.8
通讯作者:
Tseng, George C.
Tseng, George C.
中科院分区:
生物学3区
文献类型:
--
作者:
Oh, Sunghee;Kang, Dongwan D.;Tseng, George C.

文献摘要

被引文献

相似文献

动机:由于芯片上的灰尘、划痕、分辨率不足或杂交错误等缺陷,微阵列实验经常产生多个缺失值(mv)。不幸的是,许多下游算法需要一个完整的数据矩阵。这项工作的动机是确定MV归算对下游分析的影响,以及按归算精度排序的归算方法是否与归算的生物影响有很好的相关性。方法:利用8个差异表达(DE)和分类分析数据集以及8个基因聚类数据集,我们展示了缺失值imputation对统计下游分析的生物学影响,包括3种常用的DE方法、4种分类器和3种基因聚类方法。利用基于三种均方根误差(RMSE)测度的归算方法排名与基于下游分析方法的排名之间的相关性,探讨哪种RMSE测度与生物影响测度最一致,以及哪种下游分析方法对归算程序的选择最敏感。结果:DE对归算程序的选择最敏感,分类最不敏感,聚类介于两者之间。日志RMSE (LRMSE)度量与基于DE结果的归算排名的相关性最高,表明LRMSE是三个基于RMSE的度量中最具代表性的代理。在实证下游评价中,贝叶斯主成分分析和最小二乘自适应方法表现最好。
Motivation: Microarray experiments frequently produce multiple missing values (MVs) due to flaws such as dust, scratches, insufficient resolution or hybridization errors on the chips. Unfortunately, many downstream algorithms require a complete data matrix. The motivation of this work is to determine the impact of MV imputation on downstream analysis, and whether ranking of imputation methods by imputation accuracy correlates well with the biological impact of the imputation.Methods: Using eight datasets for differential expression (DE) and classification analysis and eight datasets for gene clustering, we demonstrate the biological impact of missing-value imputation on statistical downstream analyses, including three commonly employed DE methods, four classifiers and three gene-clustering methods. Correlation between the rankings of imputation methods based on three root-mean squared error (RMSE) measures and the rankings based on the downstream analysis methods was used to investigate which RMSE measure was most consistent with the biological impact measures, and which downstream analysis methods were the most sensitive to the choice of imputation procedure.Results: DE was the most sensitive to the choice of imputation procedure, while classification was the least sensitive and clustering was intermediate between the two. The logged RMSE (LRMSE) measure had the highest correlation with the imputation rankings based on the DE results, indicating that the LRMSE is the best representative surrogate among the three RMSE-based measures. Bayesian principal component analysis and least squares adaptive appeared to be the best performing methods in the empirical downstream evaluation.