Nonparametric Identification of Multivariate Mixtures

Nonparametric Identification of Multivariate Mixtures
复制标题

DOI:
--
复制
发表时间:
2010-08
期刊:
--
影响因子:
--
通讯作者:
Hiroyuki Kasahara;Katsumi Shimotsu
Hiroyuki Kasahara;Katsumi Shimotsu
中科院分区:
其他
文献类型:
--
作者:
Hiroyuki Kasahara;Katsumi Shimotsu

文献摘要

被引文献

相似文献

本文分析了k元M分量有限混合模型的可辨识性,其中每个分量分布具有独立的边缘,包括潜在类分析中的模型。在不对分量分布进行参数假设的情况下,我们研究了如何从观测数据的分布函数中识别分量的数量和分量分布。我们揭示了变量数量(K)、每个变量可以取值的数量和可识别组件的数量之间的重要联系。如果k&>=2,则分量数目(M)的下界是非参数可识别的,并且分量的最大可识别数目由每个变量取的不同值的数目确定。当M已知时,如果(I)k&gt=3,(Ii)k个变量中的两个变量取至少M个不同的值,以及(Iii)这些矩阵满足一定的秩值和特征值条件,则混合比例和分量分布是从由数据的分布函数构造的矩阵中非参数地识别的。对于未知的M情形,我们提出了一种可能从数据中识别M和分量分布的算法。我们讨论了非参数证明的一个条件及其可观测的含义。在M不能被识别的情况下,我们使用我们的识别条件来开发一个过程,通过估计由观察变量的分布函数构造的矩阵的秩来一致地估计分量数量的下界。
This article analyzes the identifiability of k-variate, M-component finite mixture models in which each component distribution has independent marginals, including models in latent class analysis. Without making parametric assumptions on the component distributions, we investigate how one can identify the number of components and the component distributions from the distribution function of the observed data. We reveal an important link between the number of variables (k), the number of values each variable can take, and the number of identifiable components. A lower bound on the number of components (M) is nonparametrically identifiable if k >= 2, and the maximum identifiable number of components is determined by the number of different values each variable takes. When M is known, the mixing proportions and the component distributions are nonparametrically identified from matrices constructed from the distribution function of the data if (i) k >= 3, (ii) two of k variables take at least M different values, and (iii) these matrices satisfy some rank and eigenvalue conditions. For the unknown M case, we propose an algorithm that possibly identifies M and the component distributions from data. We discuss a condition for nonparametric identi fication and its observable implications. In case M cannot be identified, we use our identification condition to develop a procedure that consistently estimates a lower bound on the number of components by estimating the rank of a matrix constructed from the distribution function of observed variables.