Multiple Imputation based Clustering Validation (MIV) for Big Longitudinal Trial Data with Missing Values in eHealth.

Multiple Imputation based Clustering Validation (MIV) for Big Longitudinal Trial Data with Missing Values in eHealth.
复制标题

DOI:
10.1007/s10916-016-0499-0
复制
发表时间:
2016-06
影响因子:
5.3
通讯作者:
Wang H
Wang H
中科院分区:
医学3区
文献类型:
--
作者:
Zhang Z;Fang H;Wang H

文献摘要

相似文献

网络交付的试验是电子卫生服务的一个重要组成部分。这些试验,大多是基于行为的,产生大量的异构数据,这些数据是纵向的,高维的,缺失值。无监督学习方法已广泛应用于该领域,然而,验证最优聚类数量一直具有挑战性。在我们基于多输入(MI)的模糊聚类(MIfuzzy)的基础上,我们提出了一个新的基于多输入(MI)的验证(MIV)框架和相应的MIV算法,用于聚类具有缺失值的大型纵向电子健康数据,更普遍的是基于模糊逻辑的聚类方法。具体来说,我们通过自动搜索和综合一套基于mi的验证方法和指标来检测最优聚类数量,包括常规(基于bootstrap或交叉验证)和新兴(基于模块化)的验证指标,以及用于模糊聚类的特定验证指标(Xie和Beni)。MIV性能在一个来自真实网络交付试验的大型纵向数据集上进行了演示,并使用了模拟。结果表明,基于mi的Xie和Beni模糊聚类指标更适合于此类复杂数据的最优聚类数检测。MIV概念和算法可以很容易地适应不同类型的聚类,这些聚类可以处理电子卫生服务中的大型不完整纵向试验数据。
Web-delivered trials are an important component in eHealth services. These trials, mostly behavior-based, generate big heterogeneous data that are longitudinal, high dimensional with missing values. Unsupervised learning methods have been widely applied in this area, however, validating the optimal number of clusters has been challenging. Built upon our multiple imputation (MI) based fuzzy clustering, MIfuzzy, we proposed a new multiple imputation based validation (MIV) framework and corresponding MIV algorithms for clustering big longitudinal eHealth data with missing values, more generally for fuzzy-logic based clustering methods. Specifically, we detect the optimal number of clusters by auto-searching and -synthesizing a suite of MI-based validation methods and indices, including conventional (bootstrap or cross-validation based) and emerging (modularity-based) validation indices for general clustering methods as well as the specific one (Xie and Beni) for fuzzy clustering. The MIV performance was demonstrated on a big longitudinal dataset from a real web-delivered trial and using simulation. The results indicate MI-based Xie and Beni index for fuzzy-clustering is more appropriate for detecting the optimal number of clusters for such complex data. The MIV concept and algorithms could be easily adapted to different types of clustering that could process big incomplete longitudinal trial data in eHealth services.