Clustering of Trajectories using Non-Parametric Conformal DBSCAN Algorithm

Clustering of Trajectories using Non-Parametric Conformal DBSCAN Algorithm
复制标题

DOI:
10.1109/ipsn54338.2022.00043
复制
发表时间:
2022-05
期刊:
2022 21st ACM/IEEE International Conference on Information Processing in Sensor Networks (IPSN)
影响因子:
--
通讯作者:
Haotian Wang;Jie Gao;Min-ge Xie
Haotian Wang;Jie Gao;Min-ge Xie
中科院分区:
其他
文献类型:
--
作者:
Haotian Wang;Jie Gao;Min-ge Xie

文献摘要

相似文献

技术创新为研究人类自然流动特性提供了机会。在本文中,我们研究如何在轨迹族中识别有趣的集群(通过不同的个人或其他自然定义的群体)。我们专注于粗粒度,稀疏采样的轨迹推断零星发生在无监督的设置。这是一个具有挑战性的设置,由于在选择功能和相似性措施的困难,并由于缺乏数据分布的先验知识。我们提出了一种非参数聚类算法,它对数据分布和聚类属性的先验知识做了很少的假设。我们的算法,共形DBSCAN,结合基于密度的DBSCAN聚类与统计共形预测框架。我们首先识别高度相似的轨迹组作为聚类的初始种子,类似于DBSCAN。然后,我们包括额外的轨迹属于这个集群,有保证的统计置信水平,来自一个改进的共形预测框架。这允许聚类算法自动适应不同的数据分布。我们的算法在几个人工和现实世界的数据集上的性能显着优于其他聚类算法。
Technology innovation has provided the opportunity to study the characteristics of natural human mobility. In this paper, we look at how to identify interesting clusters (by different individuals or other naturally defined groups) in a family of trajectory traces. We focus on coarse-grained, sparsely sampled trajectories inferred from sporadic occurrences in an unsupervised setting. This is a challenging setting due to difficulties in selecting features and similarity measures, and due to lack of prior knowledge of data distribution. We propose a non-parametric clustering algorithm, which makes little assumptions on prior knowledge of both data distribution and cluster properties. Our algorithm, Conformal DBSCAN, combines density-based DBSCAN clustering with the statistical conformal prediction framework. We first identify groups of highly similar trajectories as the initial seeds of clusters, similar to DBSCAN. Then we include additional trajectories that belong to this cluster, with a guaranteed statistical confidence level, derived by an improved conformal prediction framework. This allows the clustering algorithm to automatically adapt to different data distributions. Our algorithms are shown to significantly outperform alternative clustering algorithms on several artificial and real-world datasets.