ADAPTIVE ESTIMATION IN STRUCTURED FACTOR MODELS WITH APPLICATIONS TO OVERLAPPING CLUSTERING

ADAPTIVE ESTIMATION IN STRUCTURED FACTOR MODELS WITH APPLICATIONS TO OVERLAPPING CLUSTERING
复制标题

DOI:
10.1214/19-aos1877
复制
发表时间:
2020-08-01
影响因子:
4.5
通讯作者:
Wegkamp, Marten
Wegkamp, Marten
中科院分区:
数学1区
文献类型:
--
作者:
Bing, Xin;Bunea, Florentina;Wegkamp, Marten

文献摘要

被引文献

相似文献

介绍了潜在因子模型X=AZ+E中加载矩阵A的项和结构的一种新的估计方法-LOVE,因为可观测随机向量X是R-P的元素,具有相关的不可观测因子Z是R-K的元素,K未知,噪声E不相关。A的每一行都被缩放,并允许稀疏。为了识别加载矩阵A,我们需要存在纯变量,它是X的分量,通过A与一个且仅有一个潜在因素相关联。尽管因子K的个数、纯变量的个数和它们的位置都是未知的,但我们只需要对Z的协方差矩阵有一个温和的条件,并且每个潜在因子至少只有两个纯变量才能证明A是唯一定义的,直到符号排列。我们对模型可辨识性的证明是构造性的,并从X上n个观测值的样本中得到了我们的新的因子数和纯变量集的估计方法。这是我们的LOVE算法的第一步,该算法是无优化的,并且具有p(2)阶的低计算复杂性。LOVE的第二步是一个易于实现的估计A的线性规划。我们证明了所得到的估计量对于A是接近极小极大速率最优的,对于平行于(无穷)(,q)损失,对于Q>=1,直到p的对数因子,并且在许多情况下它可以是极小极大速率最优的。模型结构是由数据科学中普遍存在的重叠变量聚类问题驱动的。我们将种群水平集群定义为X的那些分量通过矩阵A与相同的不可观察到的潜在因素相关联的组,并且允许多因素关联。该算法通过纯变量分别锚定聚类,形成p维随机向量X的重叠子群。重叠聚类的潜在模型方法体现在我们的算法LOVE中;第三步,LOVE算法通过估计A的列的支持度来估计聚类,保证了零误报比例和假阴性比例控制的聚类恢复。通过对RNA-SEQ数据集的分析,说明了LOVE的实际相关性,该数据集致力于确定具有未知功能的基因的功能注释。
This work introduces a novel estimation method, called LOVE, of the entries and structure of a loading matrix A in a latent factor model X = AZ + E, for an observable random vector X is an element of R-p, with correlated unobservable factors Z is an element of R-K, with K unknown, and uncorrelated noise E. Each row of A is scaled, and allowed to be sparse. In order to identify the loading matrix A, we require the existence of pure variables, which are components of X that are associated, via A, with one and only one latent factor. Despite the fact that the number of factors K, the number of the pure variables and their location are all unknown, we only require a mild condition on the covariance matrix of Z, and a minimum of only two pure variables per latent factor to show that A is uniquely defined, up to signed permutations. Our proofs for model identifiability are constructive, and lead to our novel estimation method of the number of factors and of the set of pure variables, from a sample of size n of observations on X. This is the first step of our LOVE algorithm, which is optimization-free, and has low computational complexity of order p(2). The second step of LOVE is an easily implementable linear program that estimates A. We prove that the resulting estimator is near minimax rate optimal for A, with respect to the parallel to parallel to(infinity)(,q) loss, for q >= 1, up to logarithmic factors in p, and that it can be minimax-rate optimal in many cases of interest.The model structure is motivated by the problem of overlapping variable clustering, ubiquitous in data science. We define the population level clusters as groups of those components of X that are associated, via the matrix A, with the same unobservable latent factor, and multifactor association is allowed. Clusters are respectively anchored by the pure variables, and form overlapping subgroups of the p-dimensional random vector X. The Latent model approach to OVErlapping clustering is reflected in the name of our algorithm, LOVE.The third step of LOVE estimates the clusters from the support of the columns of the estimated A. We guarantee cluster recovery with zero false positive proportion, and with false negative proportion control. The practical relevance of LOVE is illustrated through the analysis of a RNA-seq data set, devoted to determining the functional annotation of genes with unknown function.