A self-supervised learning-based approach to clustering multivariate time-series data with missing values (SLAC-Time): An application to TBI phenotyping

A self-supervised learning-based approach to clustering multivariate time-series data with missing values (SLAC-Time): An application to TBI phenotyping
复制标题

DOI:
10.1016/j.jbi.2023.104401
复制
发表时间:
2023-02
影响因子:
4.5
通讯作者:
Hamid Ghaderi;B. Foreman;Amin Nayebi;Sindhu Tipirneni;Chandan K. Reddy;V. Subbian
Hamid Ghaderi;B. Foreman;Amin Nayebi;Sindhu Tipirneni;Chandan K. Reddy;V. Subbian
中科院分区:
医学3区
文献类型:
--
作者:
Hamid Ghaderi;B. Foreman;Amin Nayebi;Sindhu Tipirneni;Chandan K. Reddy;V. Subbian

文献摘要

被引文献

相似文献

自监督学习方法为多变量时间序列数据的聚类提供了一个很有前途的方向。然而,现实世界中的时间序列数据往往包含缺失值,现有的方法需要在聚类之前估算缺失值,这可能会导致大量的计算和噪声,并导致无效的解释。为了解决这些问题,我们提出了一种基于自我监督学习的方法来聚类具有缺失值的多变量时间序列数据(SLAC-Time)。SLAC-Time是一种基于Transformer的聚类方法,它使用时间序列预测作为代理任务,以利用未标记的数据并学习更强大的时间序列表示。该方法联合学习神经网络参数和学习表示的聚类分配。它使用K-means方法迭代地聚类学习的表示,然后利用后续的聚类分配作为伪标签来更新模型参数。为了评估我们提出的方法,我们将其应用于创伤性脑损伤(TBI)研究中的创伤性脑损伤(TBI)患者的聚类和表型分析。与TBI患者相关的临床数据通常随时间测量,并表示为以缺失值和不规则时间间隔为特征的时间序列变量。我们的实验表明,SLAC-Time优于基线K-means聚类算法的轮廓系数,Calinski Harabasz指数,Dunn指数,和Davies Bouldin指数。我们确定了三种TBI表型,它们在临床显著变量以及临床结局方面彼此不同,包括扩展格拉斯哥结局量表(GOSE)评分、重症监护室(ICU)住院时间和死亡率。实验表明,由SLAC-Time鉴定的TBI表型可潜在地用于开发靶向临床试验和治疗策略。
Self-supervised learning approaches provide a promising direction for clustering multivariate time-series data. However, real-world time-series data often include missing values, and the existing approaches require imputing missing values before clustering, which may cause extensive computations and noise and result in invalid interpretations. To address these challenges, we present aSelf-supervisedLearning-basedApproach toClustering multivariateTime-series data with missing values (SLAC-Time). SLAC-Time is a Transformer-based clustering method that uses time-series forecasting as a proxy task for leveraging unlabeled data and learning more robust time-series representations. This method jointly learns the neural network parameters and the cluster assignments of the learned representations. It iteratively clusters the learned representations with the K-means method and then utilizes the subsequent cluster assignments as pseudo-labels to update the model parameters. To evaluate our proposed approach, we applied it to clustering and phenotyping Traumatic Brain Injury (TBI) patients in the Transforming Research and Clinical Knowledge in Traumatic Brain Injury (TRACK-TBI) study. Clinical data associated with TBI patients are often measured over time and represented as time-series variables characterized by missing values and irregular time intervals. Our experiments demonstrate that SLAC-Time outperforms the baseline K-means clustering algorithm in terms of silhouette coefficient, Calinski Harabasz index, Dunn index, and Davies Bouldin index. We identified three TBI phenotypes that are distinct from one another in terms of clinically significant variables as well as clinical outcomes, including the Extended Glasgow Outcome Scale (GOSE) score, Intensive Care Unit (ICU) length of stay, and mortality rate. The experiments show that the TBI phenotypes identified by SLAC-Time can be potentially used for developing targeted clinical trials and therapeutic strategies.