K-Medoids Clustering of Data Sequences With Composite Distributions

K-Medoids Clustering of Data Sequences With Composite Distributions
复制标题

DOI:
10.1109/tsp.2019.2901370
复制
发表时间:
2018-07
影响因子:
5.4
通讯作者:
Tiexing Wang;Qunwei Li;Donald J. Bucci;Yingbin Liang;Biao Chen;P. Varshney
Tiexing Wang;Qunwei Li;Donald J. Bucci;Yingbin Liang;Biao Chen;P. Varshney
中科院分区:
工程技术1区
文献类型:
--
作者:
Tiexing Wang;Qunwei Li;Donald J. Bucci;Yingbin Liang;Biao Chen;P. Varshney

文献摘要

被引文献

相似文献

本文研究了数据序列的聚类问题,采用了k中心点算法。假设所有的数据序列都是从未知的连续分布中生成的,这些分布形成聚类,每个聚类包含一组位置接近的分布(基于分布之间的一定距离度量)。假设最大群内距离小于最小群间距离,并且两个值都是已知的。目标是将数据序列分组在一起,如果它们的底层生成分布(未知)属于一个聚类。针对分布簇个数已知和未知的情况,提出了基于分布距离度量的k中心点算法。上界的错误概率和收敛结果在大样本制度也提供。它示出的错误概率指数衰减快,在每个数据序列中的样本数趋于无穷大。当满足某些条件时,误差指数具有简单的形式,而与应用的距离度量无关。特别是,误差指数的特征在于当使用Kolmogrov-Smirnov距离或最大平均差异作为距离度量时。仿真结果验证了分析的正确性。
This paper studies clustering of data sequences using the k-medoids algorithm. All the data sequences are assumed to be generated from unknown continuous distributions, which form clusters with each cluster containing a composite set of closely located distributions (based on a certain distance metric between distributions). The maximum intracluster distance is assumed to be smaller than the minimum intercluster distance, and both values are assumed to be known. The goal is to group the data sequences together if their underlying generative distributions (which are unknown) belong to one cluster. Distribution distance metrics based k-medoids algorithms are proposed for known and unknown number of distribution clusters. Upper bounds on the error probability and convergence results in the large sample regime are also provided. It is shown that the error probability decays exponentially fast as the number of samples in each data sequence goes to infinity. The error exponent has a simple form regardless of the distance metric applied when certain conditions are satisfied. In particular, the error exponent is characterized when either the Kolmogrov–Smirnov distance or the maximum mean discrepancy are used as the distance metric. Simulation results are provided to validate the analysis.