Fuzzy partition technique for clustering Big Urban dataset

Fuzzy partition technique for clustering Big Urban dataset
复制标题

DOI:
10.1109/sai.2016.7555984
复制
发表时间:
2016-07
期刊:
2016 SAI Computing Conference (SAI)
影响因子:
--
通讯作者:
Ahmad Alshami;Weisi Guo;Ganna Pogrebna
Ahmad Alshami;Weisi Guo;Ganna Pogrebna
中科院分区:
其他
文献类型:
--
作者:
Ahmad Alshami;Weisi Guo;Ganna Pogrebna

文献摘要

相似文献

智慧城市正在从各种数据源收集和产生海量数据,如当地气象站、激光雷达数据、手机传感器、物联网(IoT)等。要将如此海量的数据用于潜在利益,使用高效和有效的大数据算法存储和分析数据至关重要。然而,由于许多挑战,这可能是有问题的。本文探讨了其中的一些挑战,并测试了两种用于集群这类大城市数据集的分区算法的性能。对K-均值和模糊c-均值(FCM)这两种简单的聚类算法进行了测试。城市数据聚类的目的是根据特定的属性将其归类为同质组。大城市数据以紧凑的格式聚类,代表了整个数据的信息,这有助于研究人员更有效地处理这些重组后的数据。为了实现这一目标,这两种技术被用来对比大量的激光雷达数据,以显示它们在相同的硬件设置上的表现。我们的实验结论是,当呈现这种类型的数据集时,FCM的性能优于K-Means,但后者对硬件利用率的要求较低。
Smart cities are collecting and producing massive amount of data from various data sources such as local weather stations, LIDAR data, mobile phones sensors, Internet of Things (IoT) etc. To use such large volume of data for potential benefits, it is important to store and analyse data using efficient and effective big data algorithms. However, this can be problematic due to many challenges. This article explores some of these challenges and tested the performance of two partition algorithms for clustering such Big Urban Datasets. Two handy clustering algorithms the K-Means vs. the Fuzzy c-Mean (FCM) were put to the test. The purpose of clustering urban data is to categorize it into homogeneous groups according to specific attributes. Clustering Big Urban Data in compact format represents the information of the whole data and this can benefit researchers to deal with this reorganised data much efficiently. To achieve this end, the two techniques were utilised against a large set of Lidar data to show how they perform on the same hardware set-up. Our experiments conclude that FCM outperformed the K-Means when presented with such type of dataset, however the latter is less demanding on the hardware utilisation.