High Performance Computing

High Performance Computing
复制标题

DOI:
10.1007/978-3-030-16205-4
复制
发表时间:
2018
期刊:
Communications in Computer and Information Science
影响因子:
--
通讯作者:
Suzanne Lun
Suzanne Lun
中科院分区:
其他
文献类型:
--
作者:
Suzanne Lun

文献摘要

被引文献

相似文献

数据驱动的医疗保健的价值在于可以检测住院护理、治疗、预防和疾病理解的新模式,或者预测住院时间、住院费用或住院期间是否可能发生死亡。从临床数据中建模精确的患者表型表示是具有挑战性的,因为其高维,噪声和缺失的数据要处理到一个新的低维空间中。同样,将无监督学习模型处理成不断增长的临床数据也会带来许多问题,比如算法复杂性,比如建模收敛时间和内存容量。本文提出了DiagnoseNET框架自动化病人表型提取,并将其应用于预测不同的医疗目标。它提供了三个高级功能:完整的工作流程编排到阶段流水线,用于挖掘临床数据并使用无监督特征表示来初始化监督模型;用于训练并行和分布式深度神经网络的数据资源管理。作为一个研究案例,我们使用了来自入院和医院服务的临床数据集来构建一个通用的住院患者表型表示,用于不同的医疗目标,第一个目标是对住院患者护理的主要目的进行分类。该研究的重点是根据数据的维度、模型复杂度、选择的工人数量和内存容量来管理数据,以便在Mini-Cluster Jetson TX 2上训练无监督的staked去噪自动编码器。因此,映射适合计算资源的任务是最小化模型收敛所需的时期数量,减少执行时间并最大化能量效率的关键因素。
The value of data-driven healthcare is the possibility to detect new patterns for inpatient care, treatment, prevention, and comprehension of disease or to predict the duration of hospitalization, its cost or whether death is likely to occur during the hospital stay. Modeling precise patients phenotype representation from clinical data is challenging over its high-dimensionality, noisy and missing data to be processed into a new low-dimensionality space. Likewise, processing unsupervised learning models into a growing clinical data raises many issues, in terms of algorithmic complexity, such as time to model convergence and memory capacity. This paper presents DiagnoseNET framework to automate patient phenotype extractions and apply them to predict different medical targets. It provides three high-level features: a full-workflow orchestration into stage pipelining for mining clinical data and using unsupervised feature representations to initialize supervised models; a data resource management for training parallel and distributed deep neural networks. As a case of study, we have used a clinical dataset from admission and hospital services to build a general purpose inpatient phenotype representation to be used in different medical targets, the first target is to classify the main purpose of inpatient care. The research focuses on managing the data according to its dimensions, the model complexity, the workers number selected and the memory capacity, for training unsupervised staked denoising auto-encoders over a Mini-Cluster Jetson TX2. Therefore, mapping tasks that fit over computational resources is a key factor to minimize the number of epochs necessary to model converge, reducing the execution time and maximizing the energy efficiency.