FINDING PATTERN BEHAVIOR IN TEMPORAL DATA USING FUZZY CLUSTERING

FINDING PATTERN BEHAVIOR IN TEMPORAL DATA USING FUZZY CLUSTERING
复制标题

使用模糊聚类查找时态数据中的模式行为

DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
K. Cheok
K. Cheok
中科院分区:
--
文献类型:
--
作者:
R. Haskell;Darrin M. Hanna;P. Li;K. Cheok

文献摘要

被引文献

相似文献

采用基于模糊等价关系的聚类技术对时态数据进行特征描述。在初始时间段内收集的数据被分成不同的集群。这些团簇的特征是它们的质心。在随后的时间段内形成的集群要么与现有集群合并,要么添加到集群列表中。得到的聚类质心列表称为聚类组,它描述了一组特定时间数据的行为特征。在随后的一段时间内形成的新集群与集群组相似的程度由相似性度量q来表征。该技术已应用于检测驾驶员行为的问题。许多不同的聚类技术被用于分析多变量数据[1]。这些方法已被应用于知识发现和数据挖掘等领域。标准聚类算法将每个数据样本分配到众多聚类中的一个,其中特定聚类中的所有样本在某种意义上是相似的。模糊聚类算法并不坚持每个样本必须只属于一个聚类,而是样本可以不同程度地属于多个聚类。最著名的模糊聚类算法是模糊c均值算法[3],该算法要求给定聚类中心的个数c。另一种不需要事先知道聚类数量的聚类方法是基于模糊等价关系的使用[4,5]。该方法形成一个模糊相容关系矩阵Q,矩阵中的每一项表示两个不同样本之间的接近程度。值为1(在主对角线上)表示样本与自身接近的程度,值为0表示数据集中被最大可能距离分隔的样本。Q的传递闭包通过对模糊集[5]选择不同的α-切,引起数据的清晰划分(产生不同数量的聚类)。本文将使用以这种方式形成的簇来描述时态数据中的模式行为。这项研究的动机是希望通过监测汽车电脑已经测量到的信号来描述驾驶员的行为
A clustering technique based on a fuzzy equivalence relation is used to characterize temporal data. Data collected during an initial time period are separated into clusters. These clusters are characterized by their centroids. Clusters formed during subsequent time periods are either merged with an existing cluster or added to the cluster list. The resulting list of cluster centroids, called a cluster group, characterizes the behavior of a particular set of temporal data. The degree to which new clusters formed in a subsequent time period are similar to the cluster group is characterized by a similarity measure, q. This technique has been applied to the problem of detecting driver behavior. INTRODUCTION Many different clustering techniques have been used for analyzing multivariate data [1]. These methods have been applied to problems in knowledge discovery and data mining [2]. The standard clustering algorithms assign each data sample to one of many clusters in which all samples in a particular cluster are similar in some sense. Fuzzy clustering algorithms do not insist that each sample must belong to only one cluster, but rather samples can belong to more than one cluster to varying degrees. The most well known fuzzy clustering algorithm is the fuzzy c-means algorithm [3] that requires that the number of cluster centers, c, be given. A different clustering approach that does not require the number of clusters to be known beforehand is based on the use of fuzzy equivalence relations [4, 5]. In this method a fuzzy compatibility relation matrix, Q, is formed in which each entry in the matrix represents the degree to which two different samples are close to each other. A value of 1 (on the main diagonal) represents the degree to which a sample is close to itself, while a value of 0 represents samples separated by the largest possible distance in the data set. The transitive closure of Q will induce crisp partitions of the data (resulting in different numbers of clusters) by choosing different α-cuts of a fuzzy set [5]. Clusters formed in this manner will be used in this paper to characterize pattern behavior in temporal data. This research was motivated by the desire to characterize a driver’s behavior by monitoring signals that are already being measured by the car’s computer