Multitask learning and benchmarking with clinical time series data

Multitask learning and benchmarking with clinical time series data
复制标题

DOI:
10.1038/s41597-019-0103-9
复制
发表时间:
2019-06-17
期刊:
影响因子:
9.8
通讯作者:
Galstyan, Aram
Galstyan, Aram
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Harutyunyan, Hrayr;Khachatrian, Hrant;Galstyan, Aram

文献摘要

被引文献

相似文献

医疗保健是数据挖掘和机器学习中最令人兴奋的边界之一。成功采用电子健康记录(EHRS)在可用于分析的数字临床数据中造成了爆炸,但是由于没有公开可用的基准数据集,因此很难衡量医疗保健研究的机器学习进展。为了解决这个问题,我们使用从重症监护室(MIMIC-III)数据库中得出的数据提出了四个临床预测基准。这些任务涵盖了一系列临床问题,包括建模死亡率的风险,预测的住院时间,检测生理下降和表型分类。我们为所有四个任务提出了强大的线性和神经基准,并评估深度监督,多任务训练和特定数据的建筑修改对神经模型性能的影响。
Health care is one of the most exciting frontiers in data mining and machine learning. Successful adoption of electronic health records (EHRs) created an explosion in digital clinical data available for analysis, but progress in machine learning for healthcare research has been difficult to measure because of the absence of publicly available benchmark data sets. To address this problem, we propose four clinical prediction benchmarks using data derived from the publicly available Medical Information Mart for Intensive Care (MIMIC-III) database. These tasks cover a range of clinical problems including modeling risk of mortality, forecasting length of stay, detecting physiologic decline, and phenotype classification. We propose strong linear and neural baselines for all four tasks and evaluate the effect of deep supervision, multitask training and data-specific architectural modifications on the performance of neural models.