Enhancing statistical power in temporal biomarker discovery through representative shapelet mining.

Enhancing statistical power in temporal biomarker discovery through representative shapelet mining.
复制标题

DOI:
10.1093/bioinformatics/btaa815
复制
发表时间:
2020-12-30
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Borgwardt K
Borgwardt K
中科院分区:
其他
文献类型:
--
作者:
Gumbsch T;Bock C;Moor M;Rieck B;Borgwardt K

文献摘要

参考文献

被引文献

相似文献

纵向数据中的时间生物标志物发现是基于检测重复出现的轨迹,即所谓的形状。搜索形状需要考虑数据中的所有连续性。虽然在以前的工作中已经减轻了多重测试的伴随问题,但检测到的形状的冗余和重叠导致先验无限数量的高度相似且结构上无意义的形状。因此,目前的时间生物标志物发现方法是不切实际的和动力不足的。我们发现,前处理或后处理的形状不足以增加的权力和实际效用。因此,我们提出了一种新的时间生物标志物发现方法:统计显著子模块子集Shapelet挖掘(S5M),检索(i)发生在数据中的短序列,(ii)与表型统计显著相关,(iii)在最大化结构多样性的同时具有可管理的数量。结构多样性是通过子模块优化修剪非代表性的形状。与模拟和真实世界数据集上的最先进方法相比,这增加了S5M的统计能力和实用性。对于进入重症监护室(ICU)的患者显示严重器官衰竭的迹象,我们发现时序模式的器官衰竭评估评分与ICU死亡率相关。 S5M是S3M的Python包中的一个选项:github.com/BorgwardtLab/S3M。
Temporal biomarker discovery in longitudinal data is based on detecting reoccurring trajectories, the so-called shapelets. The search for shapelets requires considering all subsequences in the data. While the accompanying issue of multiple testing has been mitigated in previous work, the redundancy and overlap of the detected shapelets results in an a priori unbounded number of highly similar and structurally meaningless shapelets. As a consequence, current temporal biomarker discovery methods are impractical and underpowered. We find that the pre- or post-processing of shapelets does not sufficiently increase the power and practical utility. Consequently, we present a novel method for temporal biomarker discovery: Statistically Significant Submodular Subset Shapelet Mining (S5M) that retrieves short subsequences that are (i) occurring in the data, (ii) are statistically significantly associated with the phenotype and (iii) are of manageable quantity while maximizing structural diversity. Structural diversity is achieved by pruning non-representative shapelets via submodular optimization. This increases the statistical power and utility of S5M compared to state-of-the-art approaches on simulated and real-world datasets. For patients admitted to the intensive care unit (ICU) showing signs of severe organ failure, we find temporal patterns in the sequential organ failure assessment score that are associated with in-ICU mortality. S5M is an option in the python package of S3M: github.com/BorgwardtLab/S3M.
DOI: 10.1002/prot.25461
发表时间: 2018-04
期刊: Proteins
影响因子: 2.9
作者:
Libbrecht MW;Bilmes JA;Noble WS
通讯作者: Noble WS
DOI: 10.1093/bioinformatics/bty1020
发表时间: 2019-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Llinares-Lopez, Felipe;Papaxanthos, Laetitia;Borgwardt, Karsten
通讯作者: Borgwardt, Karsten
DOI: 10.1137/1.9781611972795.41
发表时间: 2009-01-01
期刊: Proceedings of the ... SIAM International Conference on Data Mining. SIAM International Conference on Data Mining
影响因子: --
作者:
Mueen, Abdullah;Keogh, Eamonn;Westover, Brandon
通讯作者: Westover, Brandon
DOI: 10.1007/bf01588971
发表时间: 1978-01-01
影响因子: 2.7
作者:
NEMHAUSER, GL;WOLSEY, LA;FISHER, ML
通讯作者: FISHER, ML