Detecting Outliers in Data with Correlated Measures

Detecting Outliers in Data with Correlated Measures
复制标题

DOI:
10.1145/3269206.3271798
复制
发表时间:
2018-08
期刊:
Proceedings of the 27th ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Yu-Hsuan Kuo;Z. Li;Daniel Kifer
Yu-Hsuan Kuo;Z. Li;Daniel Kifer
中科院分区:
其他
文献类型:
--
作者:
Yu-Hsuan Kuo;Z. Li;Daniel Kifer

文献摘要

被引文献

相似文献

传感器技术的进步使大规模数据集的收集成为可能。这样的数据集可能会非常嘈杂,并且通常包含大量由传感器故障或人为操作故障导致的离群值。为了将这些数据用于现实世界的应用程序,检测离群值是至关重要的,这样根据这些数据集构建的模型就不会受到离群值的影响。在本文中,我们提出了一种新的离群点检测方法,该方法利用数据中的相关性(例如,出租车出行距离与出行时间)。与现有的孤立点检测方法不同,我们建立了一个稳健的回归模型,该模型显式地对孤立点进行建模,并在模型拟合的同时检测孤立点。我们在真实世界的数据集上验证我们的方法,对照专门为每个数据集设计的方法以及最先进的离群点检测器。我们的孤立点检测方法取得了更好的性能,证明了我们方法的健壮性和通用性。最后,我们报告了一些由非典型事件导致的离群值的有趣案例研究。
Advances in sensor technology have enabled the collection of large-scale datasets. Such datasets can be extremely noisy and often contain a significant amount of outliers that result from sensor malfunction or human operation faults. In order to utilize such data for real-world applications, it is critical to detect outliers so that models built from these datasets will not be skewed by outliers. In this paper, we propose a new outlier detection method that utilizes the correlations in the data (e.g., taxi trip distance vs. trip time). Different from existing outlier detection methods, we build a robust regression model that explicitly models the outliers and detects outliers simultaneously with the model fitting. We validate our approach on real-world datasets against methods specifically designed for each dataset as well as the state of the art outlier detectors. Our outlier detection method achieves better performances, demonstrating the robustness and generality of our method. Last, we report interesting case studies on some outliers that result from atypical events.