Cluster analysis and related techniques in medical research.

Cluster analysis and related techniques in medical research.
复制标题

DOI:
10.1177/096228029200100103
复制
发表时间:
1992-01-01
影响因子:
2.3
通讯作者:
McLachlan, G J
McLachlan, G J
中科院分区:
医学3区
文献类型:
--
作者:
McLachlan, G J

文献摘要

被引文献

相似文献

在本文中,我们回顾了根据临床和/或实验室类型观察对患者进行分类的聚类分析方法。考虑了分层和非分层聚类方法,尽管重点是后者,特别关注基于混合似然的方法。为了将给定数据集划分为 g 个簇,该方法使用最大似然法拟合 g 个分量的混合模型。因此,它为聚类提供了良好的统计基础。尽管理论和计算困难仍然存在,但数据中有多少簇这一重要但困难的问题可以在标准统计理论的框架内得到解决。据报道,两个案例研究分别涉及一些血友病和糖尿病数据的聚类分析,证明了基于混合可能性的聚类方法。
In this paper we review methods of cluster analysis in the context of classifying patients on the basis of clinical and/or laboratory type observations. Both hierarchical and non-hierarchical methods of clustering are considered, although the emphasis is on the latter type, with particular attention devoted to the mixture likelihood-based approach. For the purposes of dividing a given data set into g clusters, this approach fits a mixture model of g components, using the method of maximum likelihood. It thus provides a sound statistical basis for clustering. The important but difficult question of how many clusters are there in the data can be addressed within the framework of standard statistical theory, although theoretical and computational difficulties still remain. Two case studies, involving the cluster analysis of some haemophilia and diabetes data respectively, are reported to demonstrate the mixture likelihood-based approach to clustering.