GPU Implementation of Pairwise Gaussian Mixture Models for Multi-Modal Gene Co-Expression Networks

GPU Implementation of Pairwise Gaussian Mixture Models for Multi-Modal Gene Co-Expression Networks
复制标题

DOI:
10.1109/access.2019.2951284
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Ficklin, Stephen P.
Ficklin, Stephen P.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Shealy, Benjamin T.;Burns, Josh J. R.;Ficklin, Stephen P.

文献摘要

被引文献

相似文献

基因共表达网络(GCN)广泛应用于生物信息学研究,基于所有表达基因之间的成对相关性对生物体进行系统级分析。对于包含多个来源样本的大型数据集,基因对可以表现出多种共表达模式,这会混淆典型的相关方法。在计算每个模式的相关性之前,可以使用诸如高斯混合模型(GMM)的聚类方法以无监督的方式分离每个基因对的模式。然而,成对聚类显着增加了构建 GCN 的计算成本,因为必须为每个基因对评估多个聚类模型,并且基因对的数量随着基因数量的增加而快速增长。在本文中,我们提出了一种用于多模式 GCN 构建的异构、高吞吐量多 CPU/GPU 软件包,在知识独立网络构建 (KINC) 软件的第 3 版中实现。我们确定 GPU 实现的多个执行参数的最佳值,并对最多 8 个 CPU/GPU 的 CPU 和 GPU 实现进行基准测试。我们的 GPU 实现比相应的 CPU 实现实现了 167$\times$ 加速,比 KINCv1 实现了 500$\times$ 加速。
Gene co-expression networks (GCNs) are widely used in bioinformatics research to perform system-level analyses of organisms based on the pairwise correlation between all expressed genes. For large datasets which contain samples from multiple sources, gene pairs can exhibit multiple modes of co-expression which confound typical correlation approaches. A clustering method such as Gaussian Mixture Models (GMMs) may be used to separate the modes of each gene pair in an unsupervised manner, prior to computing the correlation of each mode. However, pairwise clustering significantly increases the computational cost of constructing a GCN, as several clustering models must be evaluated for each gene pair, and the number of gene pairs grows rapidly with the number of genes. In this paper, we present a heterogeneous, high-throughput multi-CPU/GPU software package for multi-modal GCN construction, implemented in version 3 of the Knowledge Independent Network Construction (KINC) software. We determine the optimal values for several execution parameters of the GPU implementation, and we benchmark our CPU and GPU implementations for up to 8 CPUs/GPUs. Our GPU implementation achieves a 167$\times$ speedup over the corresponding CPU implementation, as well as a 500$\times$ speedup over KINCv1.