An Efficient Approach to Clustering in Large Multimedia Databases with Noise

An Efficient Approach to Clustering in Large Multimedia Databases with Noise
复制标题

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
V. Guralnik;D. Wijesekera;J. Srivastava
V. Guralnik;D. Wijesekera;J. Srivastava
中科院分区:
其他
文献类型:
--
作者:
V. Guralnik;D. Wijesekera;J. Srivastava

文献摘要

被引文献

相似文献

序列数据在许多应用程序中自然出现,可以将其视为事件的排序,其中每个事件都有一个相关的发生时间。事件序列的一个重要特征是事件的发生,即以某种模式发生的事件的集合。特别有趣的是发生的事件,即发生频率高于某一阈值的事件。本文研究了序列数据中f~频繁事件的挖掘问题。我们提出了一个有效挖掘频繁事件的框架,它在许多方面超越了以前的工作。首先,我们提出一种语言来指定感兴趣的情节。其次,我们描述了一种新的数据结构,称为顺序模式树(SP树),它以一种非常紧凑的方式捕获模式语言中指定的关系。第三,我们展示了如何通过标准的自底向上挖掘算法使用该数据结构以有效的方式生成频繁的剧集。最后,我们展示了如何通过共享共同条件来优化SP树,并且只对每个这样的表达式求值一次。我们提出了对所提出的技术进行评估的结果。
Sequence data arise naturally in many applications, and can be viewed as an ordering of events, where each event has an associated time of occurrence. An important characteristic of event sequences i the occurrence of episodes, i.e. a collection of events occurring in a certain pattern. Of special interest axe ~r~uent episodes, i.e. episodes occurring with a frequency above a certain threshold. In this paper, we study the problem of mining for f~equent episodes in sequence data. We present a framework for efficient mining of frequent episodes which goes beyond previous work in a number of ways. First, we present a language for specifying episodes of interest. Second, we describe a novel data structure, called the sequential pattern tree (SP Tree), which captures the relationships specified in the pattern language in a very compact manner. Third, we show how this data structure can be used by a standard bottomup mining algorithm to generate frequent episodes in an efficient manner. Finally, we show how the SP Tree can be optimized by sharing common conditions, and evaluating each such expression only once. We present the results of an evaluation of the proposed techniques.