Bayesian-Optimization-Based Peak Searching Algorithm for Clustering in Wireless Sensor Networks

Bayesian-Optimization-Based Peak Searching Algorithm for Clustering in Wireless Sensor Networks
复制标题

DOI:
10.3390/jsan7010002
复制
发表时间:
2018-01
期刊:
J. Sens. Actuator Networks
影响因子:
--
通讯作者:
Tianyu Zhang;Qian Zhao;Kilho Shin;Y. Nakamoto
Tianyu Zhang;Qian Zhao;Kilho Shin;Y. Nakamoto
中科院分区:
其他
文献类型:
--
作者:
Tianyu Zhang;Qian Zhao;Kilho Shin;Y. Nakamoto

文献摘要

被引文献

相似文献

我们提出了一种新的峰值搜索算法(PSA),它使用贝叶斯优化来查找数据集中的概率峰值,从而提高聚类算法的速度和准确性。无线传感器网络 (WSN) 在分析和使用收集的传感数据的各种应用中变得越来越普遍。通常,收集的数据不能直接用于采用机器学习技术的现代数据分析问题,因为此类数据缺乏指定其用户目的的附加信息(例如数据标签)。当未提供附加信息时,通常会使用将数据集中的数据划分为簇的聚类算法。然而,传统的聚类算法,例如期望最大化 (EM) 和 k-means 算法,需要大量迭代才能形成聚类。因此,处理速度很慢,并且由于此类算法形成聚类的方式,聚类结果变得不太准确。 PSA 解决了这些问题,我们对其进行了调整,使其与 EM 和 k-means 算法一起使用,创建了修改后的 PSEM 和 PSK-means 算法。我们的模拟结果表明,我们提出的 P S E M 和 P S k-means 算法显着减少了所需的聚类迭代次数(1.99 至 6.3 倍),并且对于合成数据集产生的聚类精度比传统 EM 和增强型 k-means (k-means++) 算法准确 1.69 至 1.71 倍。此外,在旨在检测异常值的WSN应用模拟中,P S E M正确识别了真实数据集中的异常值,减少了约1.88倍的迭代次数,并且P S E M的准确率最多比EM高1.29倍。
We propose a new peak searching algorithm (PSA) that uses Bayesian optimization to find probability peaks in a dataset, thereby increasing the speed and accuracy of clustering algorithms. Wireless sensor networks (WSNs) are becoming increasingly common in a wide variety of applications that analyze and use collected sensing data. Typically, the collected data cannot be directly used in modern data analysis problems that adopt machine learning techniques because such data lacks additional information (such as data labels) specifying its purpose of users. Clustering algorithms that divide the data in a dataset into clusters are often used when additional information is not provided. However, traditional clustering algorithms such as expectation–maximization (EM) and k - m e a n s algorithms require massive numbers of iterations to form clusters. Processing speeds are therefore slow, and clustering results become less accurate because of the way such algorithms form clusters. The PSA addresses these problems, and we adapt it for use with the EM and k - m e a n s algorithms, creating the modified P S E M and P S k - m e a n s algorithms. Our simulation results show that our proposed P S E M and P S k - m e a n s algorithms significantly decrease the required number of clustering iterations (by 1.99 to 6.3 times), and produce clustering that, for a synthetic dataset, is 1.69 to 1.71 times more accurate than it is for traditional EM and enhanced k - m e a n s ( k - m e a n s ++) algorithms. Moreover, in a simulation of WSN applications aimed at detecting outliers, P S E M correctly identified the outliers in a real dataset, decreasing iterations by approximately 1.88 times, and P S E M was 1.29 times more accurate than EM at a maximum.