kmlShape: An Efficient Method to Cluster Longitudinal Data (Time-Series) According to Their Shapes.

kmlShape: An Efficient Method to Cluster Longitudinal Data (Time-Series) According to Their Shapes.
复制标题

DOI:
10.1371/journal.pone.0150738
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Subtil F
Subtil F
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Genolini C;Ecochard R;Benghezal M;Driss T;Andrieu S;Subtil F

文献摘要

被引文献

相似文献

纵向数据是在一段时间内反复测量每个变量的数据。分析这类数据的一种可能性是将它们聚类。大多数聚类方法将在给定时间点上轨迹接近的个体聚在一起。这些方法对局部接近的轨迹进行分组,但不一定是那些形状相似的轨迹。然而,在某些情况下,一种现象的发展过程可能比它发生的时刻更重要。因此,人们希望实现一种划分,即每个群体聚集轨迹形状相似的个体,无论它们之间的时间间隔如何。在本文中,我们提出了一种基于轨迹形状而不是经典距离的纵向数据划分算法。由于该算法耗时,我们提出了两种数据简化过程,使其适用于高维数据集。在阿尔茨海默病的应用中,该算法揭示了经典方法没有发现的“快速衰退”患者群体。在对女性月经周期的另一个应用中,该算法显示,与目前的文献相反,黄体生成素在很大比例的女性(22%)中呈现两个峰值。
Longitudinal data are data in which each variable is measured repeatedly over time. One possibility for the analysis of such data is to cluster them. The majority of clustering methods group together individual that have close trajectories at given time points. These methods group trajectories that are locally close but not necessarily those that have similar shapes. However, in several circumstances, the progress of a phenomenon may be more important than the moment at which it occurs. One would thus like to achieve a partitioning where each group gathers individuals whose trajectories have similar shapes whatever the time lag between them. In this article, we present a longitudinal data partitioning algorithm based on the shapes of the trajectories rather than on classical distances. Because this algorithm is time consuming, we propose as well two data simplification procedures that make it applicable to high dimensional datasets. In an application to Alzheimer disease, this algorithm revealed a “rapid decline” patient group that was not found by the classical methods. In another application to the feminine menstrual cycle, the algorithm showed, contrarily to the current literature, that the luteinizing hormone presents two peaks in an important proportion of women (22%).