An Efficient Technique for Mining Approximately Frequent Substring Patterns

An Efficient Technique for Mining Approximately Frequent Substring Patterns
复制标题

DOI:
10.1109/icdmw.2007.121
复制
发表时间:
2007-10
期刊:
Seventh IEEE International Conference on Data Mining Workshops (ICDMW 2007)
影响因子:
--
通讯作者:
Xiaonan Ji;J. Bailey
Xiaonan Ji;J. Bailey
中科院分区:
其他
文献类型:
--
作者:
Xiaonan Ji;J. Bailey

文献摘要

被引文献

相似文献

在广泛的应用中,顺序模式用于发现知识。然而,在许多情况下,由于短长度或低支持,模式质量可能很低。此外,对于像蛋白质这样的密集数据集,大多数顺序模式挖掘算法返回的模式数量非常大,难以处理和分析。然而,通过放宽频率的定义并允许一些不匹配,有可能发现更高质量的模式。我们将这些模式称为频繁近似子串或fas模式,并引入一种称为FAS-Miner的算法来有效地处理挖掘任务。在真实世界蛋白质和DNA数据集上的实验表明,与标准的顺序挖掘方法相比,FAS-Miner可以发现更长的长度和更高的支持度的模式。
Sequential patterns are used to discover knowledge in a wide range of applications. However, in many scenarios pattern quality can be low, due to short lengths or low supports. Furthermore, for dense datasets such as proteins, most of the sequential pattern mining algorithms return a tremendously large number of patterns, which are difficult to process and analyze. However, by relaxing the definition of frequency and allowing some mismatches, it is possible to discover higher quality patterns. We call these patterns Frequent Approximate Substrings or FAS-patterns and we introduce an algorithm called FAS-Miner, to handle the mining task efficiently. The experiments on real-world protein and DNA datasets show that FAS-Miner can discover patterns of much longer lengths and higher supports than standard sequential mining approaches.