An Efficient Technique for Mining Approximately Frequent Substring Patterns
An Efficient Technique for Mining Approximately Frequent Substring Patterns
复制标题
DOI:
10.1109/icdmw.2007.121
复制
发表时间:
2007-10
期刊:
影响因子:
--
通讯作者:
Xiaonan Ji;J. Bailey
中科院分区:
文献类型:
--
作者:
Xiaonan Ji;J. Bailey
Sequential patterns are used to discover knowledge in a wide range of applications. However, in many scenarios pattern quality can be low, due to short lengths or low supports. Furthermore, for dense datasets such as proteins, most of the sequential pattern mining algorithms return a tremendously large number of patterns, which are difficult to process and analyze. However, by relaxing the definition of frequency and allowing some mismatches, it is possible to discover higher quality patterns. We call these patterns Frequent Approximate Substrings or FAS-patterns and we introduce an algorithm called FAS-Miner, to handle the mining task efficiently. The experiments on real-world protein and DNA datasets show that FAS-Miner can discover patterns of much longer lengths and higher supports than standard sequential mining approaches.