A novel f-divergence based generative adversarial imputation method for scRNA-seq data analysis.

A novel f-divergence based generative adversarial imputation method for scRNA-seq data analysis.
复制标题

DOI:
10.1371/journal.pone.0292792
复制
发表时间:
2023
期刊:
影响因子:
3.7
通讯作者:
--
中科院分区:
综合性期刊3区
文献类型:
--
作者:

文献摘要

相似文献

对单细胞RNA测序(scRNA-seq)数据的综合分析可以增强我们对细胞多样性的理解,并有助于为个体开发个性化疗法。大量的缺失值(称为脱落)使得scRNA-seq数据的分析成为一项具有挑战性的任务。大多数传统方法对缺失值的特定分布进行了假设,这限制了它们捕获高维scRNA-seq数据的复杂性的能力。此外,传统方法的插补性能随着较高的缺失率而下降。我们提出了一种新的基于f-发散的生成对抗性插补方法,称为sc-fGAIN,用于scRNA-seq数据插补。我们的研究确定了四个f-散度函数,即交叉熵,Kullback-Leibler(KL),反向KL和Jensen-Shannon,可以有效地与生成对抗插补网络集成以生成插补值,而无需任何假设,并在数学上证明了使用sc-fGAIN算法的插补数据的分布与原始数据的分布相同。真实的scRNA-seq数据分析表明,与许多传统方法相比,sc-fGAIN算法生成的插补值具有更小的均方根误差,对不同的缺失率具有鲁棒性,并且可以降低插补变异性。f-divergence提供的灵活性允许sc-fGAIN方法适应各种类型的数据,使其成为一种更通用的方法来估算scRNA-seq数据的缺失值。
Comprehensive analysis of single-cell RNA sequencing (scRNA-seq) data can enhance our understanding of cellular diversity and aid in the development of personalized therapies for individuals. The abundance of missing values, known as dropouts, makes the analysis of scRNA-seq data a challenging task. Most traditional methods made assumptions about specific distributions for missing values, which limit their capability to capture the intricacy of high-dimensional scRNA-seq data. Moreover, the imputation performance of traditional methods decreases with higher missing rates. We propose a novel f-divergence based generative adversarial imputation method, called sc-fGAIN, for the scRNA-seq data imputation. Our studies identify four f-divergence functions, namely cross-entropy, Kullback-Leibler (KL), reverse KL, and Jensen-Shannon, that can be effectively integrated with the generative adversarial imputation network to generate imputed values without any assumptions, and mathematically prove that the distribution of imputed data using sc-fGAIN algorithm is same as the distribution of original data. Real scRNA-seq data analysis has shown that, compared to many traditional methods, the imputed values generated by sc-fGAIN algorithm have a smaller root-mean-square error, and it is robust to varying missing rates, moreover, it can reduce imputation variability. The flexibility offered by the f-divergence allows the sc-fGAIN method to accommodate various types of data, making it a more universal approach for imputing missing values of scRNA-seq data.