A GPU-accelerated algorithm for biclustering analysis and detection of condition-dependent coexpression network modules.

A GPU-accelerated algorithm for biclustering analysis and detection of condition-dependent coexpression network modules.
复制标题

DOI:
10.1038/s41598-017-04070-4
复制
发表时间:
2017-06-23
期刊:
影响因子:
4.6
通讯作者:
Cui Y
Cui Y
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Bhattacharya A;Cui Y

文献摘要

被引文献

相似文献

在分析大规模基因表达数据时,识别在一定条件下具有共同表达模式的基因组是很重要的。已经开发了许多双聚类算法来解决这个问题。然而,从大型数据集中全面发现功能一致的双聚类仍然是一个具有挑战性的问题。在这里,我们提出了一种gpu加速的双聚类算法,该算法基于搜索基因表达数据集中每个基因的最大条件相关子组(CCS)。我们将CCS与13种广泛使用的双聚类算法进行了比较。CCS在合成和真实基因表达数据集上始终优于所有13种双聚类算法。作为一种基于关联的双聚类方法,CCS还可用于寻找条件相关的共表达网络模块。我们使用C实现了CCS算法,并使用CUDA C实现了用于GPU计算的并行CCS算法。CCS的源代码可从https://github.com/abhatta3/Condition-dependent-Correlation-Subgroups-CCS获取。
In the analysis of large-scale gene expression data, it is important to identify groups of genes with common expression patterns under certain conditions. Many biclustering algorithms have been developed to address this problem. However, comprehensive discovery of functionally coherent biclusters from large datasets remains a challenging problem. Here we propose a GPU-accelerated biclustering algorithm, based on searching for the largest Condition-dependent Correlation Subgroups (CCS) for each gene in the gene expression dataset. We compared CCS with thirteen widely used biclustering algorithms. CCS consistently outperformed all the thirteen biclustering algorithms on both synthetic and real gene expression datasets. As a correlation-based biclustering method, CCS can also be used to find condition-dependent coexpression network modules. We implemented the CCS algorithm using C and implemented the parallelized CCS algorithm using CUDA C for GPU computing. The source code of CCS is available from https://github.com/abhatta3/Condition-dependent-Correlation-Subgroups-CCS.