Density-based clustering of uncertain data

Density-based clustering of uncertain data
复制标题

DOI:
10.1145/1081870.1081955
复制
发表时间:
2005-08
期刊:
--
影响因子:
--
通讯作者:
H. Kriegel;M. Pfeifle
H. Kriegel;M. Pfeifle
中科院分区:
其他
文献类型:
--
作者:
H. Kriegel;M. Pfeifle

文献摘要

被引文献

相似文献

在许多不同的应用领域中,例如传感器数据库、基于位置的服务或人脸识别系统,必须基于模糊和不确定的数据来计算气味之间的距离。通常,这些不确定对象描述之间的距离由一个数字距离值表示。基于这样的单值距离函数,标准的数据挖掘算法可以在没有任何改变的情况下工作。本文提出用距离概率函数来表示两个模糊对象之间的相似性。这些模糊距离函数为每个可能的距离值分配概率值。通过将这些模糊距离函数直接集成到数据挖掘算法中,充分利用了这些函数所提供的信息。为了证明这种一般方法的好处,我们增强了基于密度的聚类算法DBSCAN,使它可以直接对这些模糊距离函数。在基于人工和真实世界数据集的详细实验评估中,我们展示了我们新方法的特点和好处。
In many different application areas, e.g. sensor databases, location based services or face recognition systems, distances between odjects have to be computed based on vague and uncertain data. Commonly, the distances between these uncertain object descriptions are expressed by one numerical distance value. Based on such single-valued distance functions standard data mining algorithms can work without any changes. In this paper, we propose to express the similarity between two fuzzy objects by distance probability functions. These fuzzy distance functions assign a probability value to each possible distance value. By integrating these fuzzy distance functions directly into data mining algorithms, the full information provided by these functions is exploited. In order to demonstrate the benefits of this general approach, we enhance the density-based clustering algorithm DBSCAN so that it can work directly on these fuzzy distance functions. In a detailed experimental evaluation based on artificial and real-world data sets, we show the characteristics and benefits of our new approach.