COBRAC: a fast implementation of convex biclustering with compression

COBRAC: a fast implementation of convex biclustering with compression
复制标题

COBRAC:压缩凸双聚类的快速实现

DOI:
10.1093/bioinformatics/btab248
复制
发表时间:
2021
期刊:
影响因子:
5.8
通讯作者:
Chi, Eric C
Chi, Eric C
中科院分区:
生物学3区
文献类型:
--
作者:
Yi, Haidong;Huang, Le;Mishne, Gal;Chi, Eric C

文献摘要

相似文献

摘要双聚类是聚类的概括,用于识别数据矩阵的观察(行)和特征(列)中的同时分组模式。最近,双聚类任务已被表述为凸优化问题。虽然问题的这种凸面重铸具有吸引人的特性,但现有算法不能很好地扩展。为了解决这个问题并使凸双聚类成为分析更大数据的实用工具,我们提出了一种称为 COBRAC 的快速凸双聚类的实现,通过迭代压缩问题大小和解决方案路径来减少计算时间。我们将 COBRAC 应用于多个基因表达数据集,以证明其有效性和效率。除了COBRAC的独立版本外,我们还开发了相关的在线Web服务器,用于在线计算和可视化可下载的交互结果。 可用性和实现源代码和测试数据可在https://github.com/haidyi/cvxbiclustr或https://zenodo.org/record/4620218获取。 Web 服务器可在 https://cvxbiclustr.ericchi.com 上获得。补充信息补充数据可在 Bioinformaticsonline 上获得。
SummaryBiclustering is a generalization of clustering used to identify simultaneous grouping patterns in observations (rows) and features (columns) of a data matrix. Recently, the biclustering task has been formulated as a convex optimization problem. While this convex recasting of the problem has attractive properties, existing algorithms do not scale well. To address this problem and make convex biclustering a practical tool for analyzing larger data, we propose an implementation of fast convex biclustering called COBRAC to reduce the computing time by iteratively compressing problem size along with the solution path. We apply COBRAC to several gene expression datasets to demonstrate its effectiveness and efficiency. Besides the standalone version for COBRAC, we also developed a related online web server for online calculation and visualization of the downloadable interactive results.Availability and implementationThe source code and test data are available at https://github.com/haidyi/cvxbiclustr or https://zenodo.org/record/4620218. The web server is available at https://cvxbiclustr.ericchi.com.Supplementary informationSupplementary data are available atBioinformaticsonline.