Simulation-derived best practices for clustering clinical data.

Simulation-derived best practices for clustering clinical data.
复制标题

用于聚类临床数据的模拟衍生最佳实践。

DOI:
10.1016/j.jbi.2021.103788
复制
发表时间:
2021-06
影响因子:
4.5
通讯作者:
Brock, Guy
Brock, Guy
中科院分区:
医学3区
文献类型:
--
作者:
Coombes, Caitlin E.;Liu, Xin;Abrams, Zachary B.;Coombes, Kevin R.;Brock, Guy

文献摘要

参考文献

被引文献

相似文献

临床背景下的聚类分析有望提高对慢性和急性临床医学中患者表型和病程的理解。然而,仍有工作要做,以确保解决方案是严格的,有效的,可重复的。在本文中,我们评估的最佳实践相异度矩阵计算和聚类的混合类型,临床数据。我们模拟临床数据来代表临床试验,队列研究和EHR数据中的问题,包括单类型数据集(二进制,连续,分类)和4个数据混合。我们测试了5个单一的距离度量(Jaccard,汉明,高尔,曼哈顿,欧几里德)和3个混合距离度量(DAISY,Supersom和墨卡托)与3个聚类算法(层次(HC),k-medoids,自组织映射(SOM))。通过调整后的兰德指数(ARI)和轮廓宽度(SW)进行定量和可视化验证。我们将我们最好的方法应用于两个真实世界的数据集:(1)收集了247名慢性淋巴细胞白血病患者的21个特征,(2)收集了6000名重症监护室患者的40个特征。HC优于k-medoids和SOM的ARI跨数据类型。DAISY产生了最高的平均ARI混合数据类型的所有混合物,除了不平衡的混合物为主的连续数据。与其他方法相比,DAISY与HC发现了上级,可分离的集群在两个现实世界的数据集。选择适当的混合型度量允许研究者获得患者集群的最佳分离,并最大限度地利用他们的数据。混合类型数据的上级指标使用多个集中于类型的距离来处理多种数据类型。更好的疾病亚分类为靶向治疗、精准医疗、临床决策支持和改善患者预后开辟了道路。
Clustering analyses in clinical contexts hold promise to improve the understanding of patient phenotype and disease course in chronic and acute clinical medicine. However, work remains to ensure that solutions are rigorous, valid, and reproducible. In this paper, we evaluate best practices for dissimilarity matrix calculation and clustering on mixed-type, clinical data. We simulate clinical data to represent problems in clinical trials, cohort studies, and EHR data, including single-type datasets (binary, continuous, categorical) and 4 data mixtures. We test 5 single distance metrics (Jaccard, Hamming, Gower, Manhattan, Euclidean) and 3 mixed distance metrics (DAISY, Supersom, and Mercator) with 3 clustering algorithms (hierarchical (HC), k-medoids, self-organizing maps (SOM)). We quantitatively and visually validate by Adjusted Rand Index (ARI) and silhouette width (SW). We applied our best methods to two real-world data sets: (1) 21 features collected on 247 patients with chronic lymphocytic leukemia, and (2) 40 features collected on 6000 patients admitted to an intensive care unit. HC outperformed k-medoids and SOM by ARI across data types. DAISY produced the highest mean ARI for mixed data types for all mixtures except unbalanced mixtures dominated by continuous data. Compared to other methods, DAISY with HC uncovered superior, separable clusters in both real-world data sets. Selecting an appropriate mixed-type metric allows the investigator to obtain optimal separation of patient clusters and get maximum use of their data. Superior metrics for mixed-type data handle multiple data types using multiple, type-focused distances. Better subclassification of disease opens avenues for targeted treatments, precision medicine, clinical decision support, and improved patient outcomes.
DOI: 10.1136/thoraxjnl-2016-209846
发表时间: 2017-11-01
期刊: THORAX
影响因子: 10
作者:
Castaldi, Peter J.;Benet, Marta;Garcia-Aymerich, Judith
通讯作者: Garcia-Aymerich, Judith
DOI: 10.1371/journal.pone.0217696
发表时间: 2019-06-19
期刊: PLOS ONE
影响因子: 3.7
作者:
Egan, Brent M.;Sutherland, Susan E.;Sinopoli, Angelo
通讯作者: Sinopoli, Angelo
DOI: 10.1097/cin.0000000000000423
发表时间: 2018-05-01
影响因子: 1.3
作者:
Bose, Eliezer;Radhakrishnan, Kavita
通讯作者: Radhakrishnan, Kavita
DOI: 10.1109/access.2019.2903568
发表时间: 2019-01-01
期刊: IEEE ACCESS
影响因子: 3.9
作者:
Ahmad, Amir;Khan, Shehroz S.
通讯作者: Khan, Shehroz S.
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG