Comparison of clustering methods for clinical databases

Comparison of clustering methods for clinical databases
复制标题

DOI:
10.1016/j.ins.2003.03.011
复制
发表时间:
2004-02-15
影响因子:
8.1
通讯作者:
Tsumoto, S
Tsumoto, S
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hirano, S;Sun, XG;Tsumoto, S

文献摘要

被引文献

相似文献

聚类方法可以看作是对给定数据集的无监督学习。即使没有领域知识或标签(如医学专家给出的疾病名称),这些方法也会生成数据集的分区。在某些情况下,这些新生成的类导致发现一种新疾病或新概念。本文讨论了聚类方法如何在实际医疗数据集上工作。为了进行比较,在脑膜脑炎数据集上选择并评估了以下四种聚类方法:单链接聚类和完全链接聚类、Ward方法和粗糙聚类。为了进行比较,每种聚类方法都采用单一的相似性度量,即数值属性之间的马氏距离和名义属性之间的汉明距离的线性组合。从以下几个方面评价聚类方法的有效性:(1)生成的聚类质量;(2)用于生成高质量聚类的属性与临床知识之间的对应关系。实验结果表明,Ward方法选取了临床上合理的属性,得到了最好的聚类,这也表明该相似度度量方法适用于医疗数据集。(C) 2003 Elsevier Inc.版权所有。
Clustering methods can be viewed as unsupervised learning from a given dataset. Even without domain knowledge or labels such as the names of diseases given by medical experts, these methods generate partition of datasets. In some cases, these new generated classes lead to discovery of a new disease or new concept. This paper discusses how clustering methods work on a practical medical data set. For comparison, the following four clustering methods were selected and evaluated on a dataset on meningoencephalitis: single- and complete-linkage agglomerative hierarchical clustering, Ward's method and rough clustering. For comparison, a single similarity measure, a linear combination of the Mahalanobis distance between numerical attributes and the Hamming distance between nominal attributes was given to each clustering method. Usefulness of the clustering methods was evaluated from the following viewpoints: (1) the quality of generated clusters, (2) correspondence between the attributes used to generate the high-quality clusters and clinical knowledge. The experimental results showed that the best clusters were obtained using Ward's method where the clinically reasonable attributes were selected, which also suggested that this similarity measure would be applicable to the medical data sets. (C) 2003 Elsevier Inc. All rights reserved.