A Density-Based Spatial Flow Cluster Detection Method

A Density-Based Spatial Flow Cluster Detection Method
复制标题

DOI:
10.21433/b3118mf4r9rw
复制
发表时间:
2016
期刊:
Transportation research procedia
影响因子:
--
通讯作者:
Ran Tao;J. Thill
Ran Tao;J. Thill
中科院分区:
其他
文献类型:
--
作者:
Ran Tao;J. Thill

文献摘要

被引文献

相似文献

GIScience 2016短文论文集基于密度的空间流簇检测方法Ran Tao 1,Jean-Claude Thill 1部。北卡罗来纳大学夏洛特分校地理与地球科学系,9201大学城大道,北卡罗来纳州夏洛特市,电子邮件:{rtao2;jfthill}@uncc.edu摘要了解空间来源-目的地流动数据的模式和动力学一直是空间科学家的长期目标。本文提出了一种针对离散空间流数据的基于密度的聚类检测方法。其基本思想是首先在考虑端点坐标和流量长度的情况下测量流量密度,然后将其与最新的基于密度的聚类方法相结合。我们使用精心设计的合成数据集进行实验。实验结果表明,该方法能够有效地从包含不同流动密度、长度、层次的各种情况下提取流团,同时避免了流端点的可修改面积单位问题(MAUP)、空间信息丢失和短流上的误判。1.导论空间流,也称为地理参考地之间的空间相互作用,已成为众多研究领域中经久不衰的研究对象。随着位置感知技术的广泛采用和地理信息系统(GIS)的全球传播,空间交互数据在几个方面得到了丰富,包括数量、类型、可用性、无处不在和时空粒度(Yan和Thill,2009;Guo等人)。2012年)。虽然它带来了前所未有的机会来提高我们对SI过程的理解,从而丰富SI理论,但它也带来了分析挑战,即开发更多针对SI数据量身定做的数据驱动方法(Yan和Thill,2009)。作为一种常见的数据挖掘技术,聚类检测在大型空间流的探索性分析中被证明是有用的。一种方法在将来源和目的地合并之前,分别测量它们之间的空间关系,以此作为聚集流量的基础。在这里,空间关系可以是起源地或目的地区域的邻接或邻近(Guo,2009;朱和Guo,2014)。然而,这些方法对流端点的不均匀分布和特殊分区定义很敏感,而且在短距离交互时容易产生误判。另一种类型的方法使用流几何来捆绑附近的方法(Cui等人。2008年)。虽然结果通常具有理想的视觉清晰度,但这些方法会因丢失有价值的空间信息而折衷。本文介绍了一种新的方法,该方法不仅可以从不同的流动密度、长度、层次等不同的情况下提取空间流团,而且避免了MAUP、误报和信息丢失等问题。2.在各种聚类方法的方法论中,我们选择了基于密度的传统流聚类方法,这是因为它能够发现任意形状的簇和滤除噪声。此外,光学等基于密度的方法(Ankerst et al.1999)可以有效地揭示数据中的层次结构,因为它的副产品可达性图可以转换为树状图(Sander等人。2003年;Campello等人。2013年)。此后,我们首先介绍了针对空间流定制的邻近度度量;然后逐步解释了聚类方法。
GIScience 2016 Short Paper Proceedings A Density-Based Spatial Flow Cluster Detection Method Ran Tao 1 , Jean-Claude Thill 1 Dept. of Geography & Earth Sciences, University of North Carolina at Charlotte, 9201 University City Blvd, Charlotte, NC Email:{ rtao2; jfthill}@uncc.edu Abstract Understanding the patterns and dynamics of spatial origin-destination flow data has been a long- standing goal of spatial scientists. In this paper we introduce a density-based cluster detection method tailored for disaggregated spatial flow data. The basic idea is to first measure flow density considering both endpoint coordinates and flow lengths, and combine it with state-of-art density-based clustering methods. We experiment with a carefully designed synthetic dataset. The results prove that our method can effectively extract flow clusters from various situations encompassing varied flow densities, lengths, hierarchies and, at the same time, avoid issues of Modifiable Areal Unit Problem (MAUP) of flows endpoints, loss of spatial information, and false positive errors on short flows. 1. Introduction Spatial flows, also known as spatial interactions (SI) between georeferenced places, have been an enduring study object in a wide range of research fields. With the widespread adoption of location-aware technologies and the global diffusion of geographic information systems (GIS), spatial interaction data have been enriched in several respects including volume, type, availability, ubiquity, and spatiotemporal granularity (Yan and Thill 2009; Guo et al. 2012). While it brings unprecedented opportunities to improve our understanding of SI processes and thus enriching SI theories, it also brings the analytical challenge of developing more data-drive approaches tailored for SI data (Yan and Thill 2009). As a common data mining technique, cluster detection has proved useful in exploratory analysis of large sets of spatial flows. One approach measures the spatial relationships among origins and destinations, respectively, before combining them, as the basis for clustering flows. Here, spatial relationships can be contiguity or proximity of origin or destination regions (Guo 2009; Zhu and Guo 2014). However these methods are sensitive to uneven distribution and ad hoc zoning definition of flow endpoints; besides they are prone to false positive errors on short- distance interactions. Another type of methods use flow geometry to bundle nearby ones (Cui et al. 2008). While the results usually have desirable visual clarity, these methods compromise through loss of valuable spatial information. In this paper we introduce a new method that not only can extract spatial flow clusters from various situations including varying flow densities, lengths, hierarchies, but also avoids problems like MAUP, false positive errors, and loss of information. 2. Methodology Of various clustering methods, we choose to design our flow clustering method in the density- based tradition because of its capability to discover clusters of arbitrary shape and to filter out noise. Moreover, density-based methods like OPTICS (Ankerst et al. 1999) can effectively reveal hierarchical structures in the data since its byproduct, the reachability plot, is convertible to a dendrogram (Sander et al. 2003; Campello et al. 2013). Hereafter, we first introduce the proximity metric tailored to spatial flows; then we explain the clustering method step by step.