Distribution based nearest neighbor imputation for truncated high dimensional data with applications to pre-clinical and clinical metabolomics studies.

Distribution based nearest neighbor imputation for truncated high dimensional data with applications to pre-clinical and clinical metabolomics studies.
复制标题

DOI:
10.1186/s12859-017-1547-6
复制
发表时间:
2017-02-20
期刊:
影响因子:
3
通讯作者:
Brock GN
Brock GN
中科院分区:
生物学4区
文献类型:
--
作者:
Shah JS;Rai SN;DeFilippis AP;Hill BG;Bhatnagar A;Brock GN

文献摘要

被引文献

相似文献

高通量代谢组学技术使得生物样品中多种代谢物的相对丰度测量成为可能,这对生物医学研究的许多领域都是有用的。然而,代谢组学数据集中的缺失值(MV)是常见的,可能是由于技术和生物学原因而出现的。通常,这些MV由最小值代替,这可能导致下游分析中的不同结果。在这里,我们提出了K最近邻(KNN)方法的修改版本,该方法考虑了最小值处的截断,即,KNN截断(KNN-TN)。我们比较了基于KNN-TN的插补结果与其他KNN方法的结果,如基于相关性的KNN(KNN-CR)和基于欧氏距离的KNN(KNN-EU)。我们的方法假设数据遵循截断正态分布,截断点在检测限(LOD)。通过均方根误差(RMSE)测量以及代谢物列表一致性指数(MLCI)分析每种方法的有效性对下游统计检验的影响。通过广泛的模拟研究和应用程序的三个真实的数据集,我们表明,KNN-TN具有较低的RMSE值相比,其他两个KNN程序以及更简单的插补方法的基础上取代缺失值与代谢物的平均值,零值,或LOD。KNN-TN和KNN-EU的MLCI值大致相当,在大多数情况下上级其他四种方法。我们的研究结果表明,与KNN-CR和KNN-EU相比,KNN-TN在填补不同数据集的缺失值方面通常具有更好的性能,因为随机缺失与LOD相结合。本研究中显示的结果属于代谢组学领域,但该方法可适用于因LOD而缺失的任何高通量技术。本文的在线版本(doi:10.1186/s12859-017-1547-6)包含补充材料,可供授权用户使用。
High throughput metabolomics makes it possible to measure the relative abundances of numerous metabolites in biological samples, which is useful to many areas of biomedical research. However, missing values (MVs) in metabolomics datasets are common and can arise due to both technical and biological reasons. Typically, such MVs are substituted by a minimum value, which may lead to different results in downstream analyses. Here we present a modified version of the K-nearest neighbor (KNN) approach which accounts for truncation at the minimum value, i.e., KNN truncation (KNN-TN). We compare imputation results based on KNN-TN with results from other KNN approaches such as KNN based on correlation (KNN-CR) and KNN based on Euclidean distance (KNN-EU). Our approach assumes that the data follow a truncated normal distribution with the truncation point at the detection limit (LOD). The effectiveness of each approach was analyzed by the root mean square error (RMSE) measure as well as the metabolite list concordance index (MLCI) for influence on downstream statistical testing. Through extensive simulation studies and application to three real data sets, we show that KNN-TN has lower RMSE values compared to the other two KNN procedures as well as simpler imputation methods based on substituting missing values with the metabolite mean, zero values, or the LOD. MLCI values between KNN-TN and KNN-EU were roughly equivalent, and superior to the other four methods in most cases. Our findings demonstrate that KNN-TN generally has improved performance in imputing the missing values of the different datasets compared to KNN-CR and KNN-EU when there is missingness due to missing at random combined with an LOD. The results shown in this study are in the field of metabolomics but this method could be applicable with any high throughput technology which has missing due to LOD. The online version of this article (doi:10.1186/s12859-017-1547-6) contains supplementary material, which is available to authorized users.