SPIKE-AND-SLAB LASSO BICLUSTERING

SPIKE-AND-SLAB LASSO BICLUSTERING
复制标题

DOI:
10.1214/20-aoas1385
复制
发表时间:
2021-03-01
影响因子:
1.8
通讯作者:
George, Edward, I
George, Edward, I
中科院分区:
数学4区
文献类型:
--
作者:
Moran, Gemma E.;Rockova, Veronika;George, Edward, I

文献摘要

被引文献

相似文献

双聚类方法同时对样本及其相关特征进行分组。这样,双聚类方法不同于传统的聚类方法,传统的聚类方法利用整个特征集来区分样本组。推动双聚类应用的应用包括基因组数据,其目标是根据患者或样本的基因表达谱对其进行分类;以及推荐系统,它寻求根据客户的产品偏好对客户进行分组。感兴趣的双簇通常表现为数据矩阵的秩为1的子矩阵。这个子矩阵检测问题可以看作是一个因子分析问题,其中的因子和载荷都是稀疏的。本文利用Roekova和George(J.Amer)提出的钉板套索方法,提出了一种新的双聚类方法--钉板套索双聚类(SSLB)。统计学家。阿索克。113(2018)431-444)来寻找数据矩阵的这种稀疏因式分解。SSLB还采用了印度自助餐流程,然后自动选择双色菜的数量。许多双色化方法对潜在的双色体的大小做出假设;或者假设所有双色体的大小相同,或者假设双色体非常大或非常小。相比之下,SSLB可以适应寻找具有大小连续体的双星体。SSLB通过一种快速EM算法和变分步长来实现。在各种模拟设置中,SSLB的性能优于其他双聚类方法。我们将SSLB应用于微阵列数据集和单细胞RNA测序数据集,并强调SSLB可以恢复数据中具有生物学意义的结构。SSLB软件以R/C++包的形式在https://github.com/gemoran/SSLB.上提供
Biclustering methods simultaneously group samples and their associated features. In this way, biclustering methods differ from traditional clustering methods, which utilize the entire set of features to distinguish groups of samples. Motivating applications for biclustering include genomics data, where the goal is to cluster patients or samples by their gene expression profiles; and recommender systems, which seek to group customers based on their product preferences. Biclusters of interest often manifest as rank-1 submatrices of the data matrix. This submatrix detection problem can be viewed as a factor analysis problem in which both the factors and loadings are sparse. In this paper, we propose a new biclustering method called Spike-and-Slab Lasso Biclustering (SSLB) which utilizes the Spike-and-Slab Lasso of Roekova and George (J. Amer. Statist. Assoc. 113 (2018) 431-444) to find such a sparse factorization of the data matrix. SSLB also incorporates an Indian Buffet Process prior to automatically choose the number of biclusters. Many biclustering methods make assumptions about the size of the latent biclusters; either assuming that the biclusters are all of the same size, or that the biclusters are very large or very small. In contrast, SSLB can adapt to find biclusters which have a continuum of sizes. SSLB is implemented via a fast EM algorithm with a variational step. In a variety of simulation settings, SSLB outperforms other biclustering methods. We apply SSLB to both a microarray dataset and a single-cell RNA-sequencing dataset and highlight that SSLB can recover biologically meaningful structures in the data. The SSLB software is available as an R/C++ package at https://github.com/gemoran/SSLB.