Estimation of Sparse Directed Acyclic Graphs for Multivariate Counts Data

Estimation of Sparse Directed Acyclic Graphs for Multivariate Counts Data
复制标题

DOI:
10.1111/biom.12467
复制
发表时间:
2016-09-01
期刊:
影响因子:
1.9
通讯作者:
Zhong, Hua
Zhong, Hua
中科院分区:
数学3区
文献类型:
--
作者:
Han, Sung Won;Zhong, Hua

文献摘要

被引文献

相似文献

下一代测序数据,称为高通量测序数据,记录为计数数据,一般远离正态分布。在计数数据服从Poisson对数正态分布的假设下,本文提出了一个L-1惩罚似然框架和一个有效的搜索算法来估计多变量计数数据的稀疏有向无环图(DAG)的结构.在寻找解的过程中,我们使用迭代优化程序来估计潜变量的邻接矩阵和方差矩阵。仿真结果表明,我们提出的方法优于假设多元正态分布的方法,对数变换的方法。实验还表明,在稀疏网络或枢纽网络结构下,该方法的性能优于基于秩的PC方法。作为一个真实的数据例子,我们证明了所提出的方法在估计卵巢癌研究的基因调控网络的效率。
The next-generation sequencing data, called high-throughput sequencing data, are recorded as count data, which are generally far from normal distribution. Under the assumption that the count data follow the Poisson log-normal distribution, this article provides an L-1-penalized likelihood framework and an efficient search algorithm to estimate the structure of sparse directed acyclic graphs (DAGs) for multivariate counts data. In searching for the solution, we use iterative optimization procedures to estimate the adjacency matrix and the variance matrix of the latent variables. The simulation result shows that our proposed method outperforms the approach which assumes multivariate normal distributions, and the log-transformation approach. It also shows that the proposed method outperforms the rank-based PC method under sparse network or hub network structures. As a real data example, we demonstrate the efficiency of the proposed method in estimating the gene regulatory networks of the ovarian cancer study.