Learning structure and concepts in data through data clustering

Learning structure and concepts in data through data clustering
复制标题

通过数据聚类学习数据中的结构和概念

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
C. Elkan
C. Elkan
中科院分区:
--
文献类型:
--
作者:
Greg Hamerly;C. Elkan

文献摘要

被引文献

相似文献

数据聚类是机器学习的一个重要的面向应用的分支。它的目标是在没有训练信号的情况下估计一组数据的结构或密度。有许多方法来进行数据聚类,由于这些算法具有广泛的应用,因此它们的复杂性和有效性各不相同。由于人类想要分析的数据量的爆炸性增长,快速(例如线性时间)算法是必要的,但它们通常会给出质量差的结果。 在保持快速算法的运行时特性的同时,我们展示了两种改进聚类算法的修改。第一个重点是为固定数量的集群找到更好的解决方案。我们将算法分解为基本部分,并分析了这些部分如何影响聚类解决方案的质量。第二个重点是使用统计假设检验有效地估计集群的数量,以及如何以新的方式应用。 我们还讨论了应用数据聚类的任务学习的计算机程序的结构。我们展示了如何聚类可用于提高计算机处理器模拟的准确性,同时提高其效率。
Data clustering is an important and applications-oriented branch of machine learning. Its goal is to estimate the structure or density of a set of data without a training signal. There are many approaches to data clustering that vary in their complexity and effectiveness, due to the wide number of applications that these algorithms have. Due to the explosive growth of the amount of data that humans want to analyze, fast (e.g. linear-time) algorithms are necessary, but they can often give poor quality results. While maintaining the runtime characteristics of the fast algorithms, we show modifications that improve clustering algorithms in two ways. The first focus is on finding better solutions for a fixed number of clusters. We decompose the algorithms into fundamental parts, and analyze how the parts affect the quality of clustering solutions. The second focus is on estimating the number of clusters efficiently using statistical hypothesis tests, and how that may be applied in novel ways. We also discuss the application of data clustering to the task of learning the structure of computer programs. We show how clustering may be used to improve the accuracy of computer processor simulations while simultaneously improving their efficiency.