A Clustering Method Based on Rough Sets and Its Application to Knowledge Discovery in the Medical Database

A Clustering Method Based on Rough Sets and Its Application to Knowledge Discovery in the Medical Database
复制标题

基于粗糙集的聚类方法及其在医学数据库知识发现中的应用

DOI:
10.3233/978-1-60750-928-8-206
复制
发表时间:
2001
影响因子:
--
通讯作者:
Y. Hata
Y. Hata
中科院分区:
--
文献类型:
--
作者:
S. Hirano;S. Tsumoto;Tomohiro Okuzaki;Y. Hata

文献摘要

被引文献

相似文献

本文提出了一种基于粗糙集的名义数据和数值数据的聚类方法及其在医学数据库知识发现中的应用。根据基于对象之间的相对相似性定义的不可区分关系来进行分类。相似性被定义为两种相似性度量的组合:名义属性的汉明距离和数值属性的马哈拉诺比斯距离。通过将相似的等价关系修改为相同的等价关系,抑制小范畴的过度生成。对脑膜脑炎诊断数据库进行分析以验证该方法。结果表明,该方法能够很好地处理这两类属性,并发现诊断的主要因素。
This paper proposes a clustering method for nominal and numerical data based on Rough Sets and its application to knowledge discovery in the medical database. Classification is performed according to the indiscernibility relations defined on the basis of relative similarity between objects. The similarity is defined as a combination of two types of similarity measures: the Hamming distance for nominal attributes and the Mahalanobis distance for numerical attributes. Excessive generation of small category is suppressed by modifying similar equivalence relations into the same equivalence relation. An analysis of the meningoencephalitis diagnosis database was performed to validate this method. The result showed that this method could deal well with both types of attributes and discover the primary factors for diagnosis.