Compression-based data mining of sequential data

Compression-based data mining of sequential data
复制标题

DOI:
10.1007/s10618-006-0049-3
复制
发表时间:
2007-02-01
影响因子:
4.8
通讯作者:
Handley, John
Handley, John
中科院分区:
计算机科学3区
文献类型:
--
作者:
Keogh, Eamonn;Lonardi, Stefano;Handley, John

文献摘要

被引文献

相似文献

绝大多数数据挖掘算法都需要设置很多输入参数。使用充满参数的算法有两重危险。首先,不正确的设置可能会导致算法无法找到真实的模式。其次,一个可能更隐蔽的问题是,算法可能会报告并不真正存在的虚假模式,或者极大地高估了报告模式的重要性。当用户无法理解参数在数据挖掘过程中的作用时,这种情况尤其可能发生。数据挖掘算法应该具有尽可能少的参数。参数光算法将限制我们将偏见、期望和假设强加于手头问题的能力,并将让数据本身与我们说话。在这项工作中,我们展示了生物信息学、学习和计算理论的最新结果为参数轻数据挖掘范例带来了巨大的希望。这些结果与柯尔莫戈洛夫复杂性理论有很强的联系。然而,实际上,它们可以使用任何现成的压缩算法来实现,只需添加十几行代码。我们将通过在时间序列/DNA/文本/XML/视频数据集上的实证测试,证明该方法在异常/兴趣度检测、分类和聚类方面与许多最先进的方法相竞争或更好。为了进一步证明我们的方法的优势,我们将展示它在推荐印刷服务和产品时解决实际分类问题的有效性。
The vast majority of data mining algorithms require the setting of many input parameters. The dangers of working with parameter-laden algorithms are twofold. First, incorrect settings may cause an algorithm to fail in finding the true patterns. Second, a perhaps more insidious problem is that the algorithm may report spurious patterns that do not really exist, or greatly overestimate the significance of the reported patterns. This is especially likely when the user fails to understand the role of parameters in the data mining process. Data mining algorithms should have as few parameters as possible. A parameter-light algorithm would limit our ability to impose our prejudices, expectations, and presumptions on the problem at hand, and would let the data itself speak to us. In this work, we show that recent results in bioinformatics, learning, and computational theory hold great promise for a parameter-light data-mining paradigm. The results are strongly connected to Kolmogorov complexity theory. However, as a practical matter, they can be implemented using any off-the-shelf compression algorithm with the addition of just a dozen lines of code. We will show that this approach is competitive or superior to many of the state-of-the-art approaches in anomaly/interestingness detection, classification, and clustering with empirical tests on time series/DNA/text/XML/video datasets. As a further evidence of the advantages of our method, we will demonstrate its effectiveness to solve a real world classification problem in recommending printing services and products.