Modeling Sparse Data Using MLE with Applications to Microbiome Data

Modeling Sparse Data Using MLE with Applications to Microbiome Data
复制标题

DOI:
10.1007/s42519-021-00230-y
复制
发表时间:
2022-03-01
影响因子:
0.6
通讯作者:
Yang, Jie
Yang, Jie
中科院分区:
其他
文献类型:
--
作者:
Aldirawi, Hani;Yang, Jie

文献摘要

被引文献

相似文献

对诸如微生物组和转录组学(RNA-seq)数据之类的稀疏数据进行建模是非常具有挑战性的,因为零的数量过多并且分布偏态。许多概率模型已被用于建模稀疏数据,包括泊松,负二项,零膨胀泊松和零膨胀负二项模型。识别零膨胀或障碍模型的最合适的概率模型的一种方法是基于Kolmogorov-Smirnov检验的p值。识别概率模型的主要挑战是模型参数在实践中通常是未知的。本文给出了一类零膨胀模型和栅栏模型的极大似然估计。我们还推导出相应的Fisher信息矩阵,以探索估计量的渐近性质。我们包括新的概率模型,如零膨胀贝塔二项式和零膨胀贝塔负二项式模型。我们对微生物组数据的应用表明,我们的新模型比文献中常用的模型更适合于微生物组数据建模。
Modeling sparse data such as microbiome and transcriptomics (RNA-seq) data is very challenging due to the exceeded number of zeros and skewness of the distribution. Many probabilistic models have been used for modeling sparse data, including Poisson, negative binomial, zero-inflated Poisson, and zero-inflated negative binomial models. One way to identify the most appropriate probabilistic models for zero-inflated or hurdle models is based on the p-value of the Kolmogorov-Smirnov test. The main challenge for identifying the probabilistic model is that the model parameters are typically unknown in practice. This paper derives the maximum likelihood estimator for a general class of zero-inflated and hurdle models. We also derive the corresponding Fisher information matrices for exploring the estimator's asymptotic properties. We include new probabilistic models such as zero-inflated beta binomial and zero-inflated beta negative binomial models. Our application to microbiome data shows that our new models are more appropriate for modeling microbiome data than commonly used models in the literature.