Leveraging Heterogeneous Data Across International Borders in a Privacy Preserving Manner for Clinical Deep Learning
Leveraging Heterogeneous Data Across International Borders in a Privacy Preserving Manner for Clinical Deep Learning
批准号:
1822378
负责人:
Gari Clifford
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-03-15 至 2021-02-28
中文摘要
人们越来越意识到多中心临床数据库和多机构医疗数据分析的必要性,以确保研究结果的重复性和普适性。单实例数据库算法容易出现三个明显的问题。首先,在大数据科学的背景下,数据的大小与变量的数量相比,很难在不过度拟合的情况下开发复杂的预测值,而更传统的学习算法可能会导致模型过于简化,无法捕捉不同类型的医疗信息之间的重要相关影响或相互作用。其次,在单一数据库上训练和测试预测模型可能会导致学习噪音或其他不相关的当地做法或定义上的差异,这些定义与所讨论的结果相关,但没有因果关系。这导致了在其他机构或未来当实践或环境发生变化时不起作用的模式。第三,由于信任、法律问题、隐私问题和国家政策,在机构之间共享数据,特别是跨境共享数据,是非常有问题的。解决这些问题的意义有三个方面:1)它将允许创建强大的通用数据科学模型,这些模型利用来自世界各地的海量数据;2)它还将允许识别罕见疾病或患者类型,随着我们编辑数据库,这些疾病或患者类型将变得不那么罕见;以及3)可能最重要的是,它将允许自由交换数据科学模型和解决云中医疗问题的通用方法。该项目旨在开发一套分布式深度学习和云计算技术,用于跨机构和跨边界的健康和医疗数据机器学习,而不需要受保护的健康信息离开生成机构。我们的目标是创建演示程序,说明其可行性并将该架构开源。该项目的范围包括多个机构可能希望应用于其云中医疗保健数据的广泛的基于机器学习的任务集,以及围绕跨域知识转移学习的技术问题(例如,机构/人口统计)和任务(例如,分类和预测问题的类型)。该项目有三个具体目标:1)开发一个基于云的基础设施,它保留了数据的区域自治性,但允许区域之间共享部分训练的深度神经网络的参数(包括权重和超参数),以便跨域和跨任务转移学习;2)为医疗应用中的深度学习方法开发一个标准化的编码模型;以及3)评估跨多中心和跨国界的训练和测试模型的效果,方法是将性能的改善与不损失隐私保护的跨机构训练进行比较,使用敏感度、特异度、阳性预测值、接收者操作特征(ROC)曲线下的面积和模型校准等指标。目标1-3将通过获取四个数据库(包括重症监护病房脓毒症患者数据库、护理进展笔记的免费文本语料库、从经典用于说话人识别的公共语料库获取的语音记录,以及用于面部表情分类的公共全脸图像数据库)并将它们放置在不同地缘政治位置(即美国和欧洲)的云(Google、AWS和Azure)中实现,并开发一个分布式深度学习架构,该架构学习通过跨境共享权重来提高其性能,但不是敏感的患者数据。该项目有可能为该领域做出几项贡献。首先,它将证明,跨越地缘政治边界的医疗数据可以以可互操作的方式(使用FHIR标准)提供,并可用于以保护隐私的方式训练深度学习算法,从而解决健康保险、可携带性和隐私法(HIPPA)和互操作性的问题。其次,它将为几个医疗数据集和数据类型提供开源的深度学习算法,可以跨机构使用,通过一些微调(例如,通过迁移学习)来解决类似的问题。第三,它将为迁移学习(跨域和任务)提供一套开源的元算法,这些算法在容器(Dockers)中在云中实施,可以下载供本地使用或跨不同的云供应商传输。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
There is a growing awareness of the need for multi-center clinical databases and multi-institutional analyses of healthcare data to ensure reproducibility and generalizability of research findings. Single-instance database algorithms are prone to three distinct problems. First, in the context of Big Data science, the size of the data compared to the number of variables makes it difficult to develop complex predictors without overfitting, and more traditional learning algorithms may lead to over-simplified models that do not capture important related influences or interactions between different types of healthcare information. Second, training and testing predictive models on a single database can lead to learning noise or other irrelevant local practices or differences in definitions that are correlated with, but not causally related to, the outcome in question. This leads to models that do not work in other institutions or in the future when practices or the environment changes. Third, sharing data between institutions, and in particular, across borders, is extremely problematic because of trust, legal issues, privacy issues and national policies. The significance of solving these issues is threefold: 1) it would allow the creating of strong generalizable data science models, which leverage enormous pools of data from around the world; 2) it would also allow the identification of rare diseases or patient types, which, as we compile databases, become less rare; and 3) perhaps most importantly, it would allow the free exchange of data science models and generalized approaches to solving medical problems in the cloud.This project aims to develop a set of distributed deep learning and cloud computation techniques for cross-institution and cross-border machine learning on health and medical data without the need for protected health information to leave the generating institution. The goals are to create demonstration programs which illustrate feasibility and open source the architecture. The scope of this project encompasses the broad set of machine learning-based tasks multiple institutions may want to apply to their healthcare data in the cloud, as well as the technical issues surrounding transfer learning of knowledge across domains (e.g., institutions/demographics) and tasks (e.g., types of classification and prediction problems). The project has three specific aims: 1) develop a cloud-based infrastructure which preserves regional autonomy of data, but allows the sharing of parameters of the partially trained deep neural network (including weights and hyperparameters) between regions, to allow transfer learning across domains and tasks; 2) develop a standardized coded model for deep learning approaches in medical applications; and 3) evaluate the effect of training and testing the model across multiple centers and national boundaries, by comparing improvement in performance with cross-institutional training without loss of privacy protection, using metrics of sensitivity, specificity, positive predictive value, area under the receiver operating characteristic (ROC) curve and model calibration. Aims 1-3 will be achieved by taking four databases (including, a database of intensive care unit patients with sepsis, a free text corpus of nursing progress notes, voice recordings taken from a public corpus classically used for speaker identification, and a public database of full-face images used for classification of facial expressions) and placing them in the cloud (Google, AWS and Azure) at different geopolitical locations (namely US and Europe) and developing a distributed deep learning architecture that learns to improve its performance by sharing weights across borders, but not sensitive patient data. This project has the potential to make several contributions to the field. First, it will demonstrate that medical data across geopolitical boundaries can be made available in an interoperable manner (using the FHIR standard) and can be used for training of deep learning algorithms in a privacy-preserving manner, thus addressing both the concerns of Health Insurance, Portability and Privacy Act (HIPPA) and interoperability. Secondly, it will provide open-source deep learning algorithms for several medical datasets and data types that can be used across institutions to solve similar problems with some fine-tuning (e.g., via transfer learning). Third, it will provide a set of open-source meta algorithms for transfer learning (across domains and tasks) implemented on the cloud in containers (dockers) that can be downloaded for local use or transferred across the different cloud vendors.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1097/ccm.0000000000004145
发表时间:
2020-02-01
期刊:
CRITICAL CARE MEDICINE
影响因子:
8.8
作者:
[Reyna, Matthew A., Josef, Christopher S., Sharma, Ashish]
通讯作者:
Sharma, Ashish
DOI:
10.1145/3365211
发表时间:
2020-02
期刊:
ACM Transactions on Information Systems (TOIS)
影响因子:
--
作者:
[Faizan Ahmad;A. Abbasi;Jingjing Li;David G. Dobolyi;Richard G. Netemeyer;G. Clifford;Hsinchun Chen]
通讯作者:
Faizan Ahmad;A. Abbasi;Jingjing Li;David G. Dobolyi;Richard G. Netemeyer;G. Clifford;Hsinchun Chen
DeepAISE on FHIR — An Interoperable Real-Time Predictive Analytic Platform for Early Prediction of Sepsis
FHIR 上的 DeepAISE — 用于脓毒症早期预测的可互操作实时预测分析平台
DOI:
--
发表时间:
2018
期刊:
AMIA Annual Symposium proceedings
影响因子:
--
作者:
[Lakshman, Vidyashankar, Amrollahi, Fatemeh, Koppisetty, Veera Supraja, Shashikumar, Supreeth P., Sharma, Ashish, Nemati, Shamim]
通讯作者:
Nemati, Shamim
BD Spokes: SPOKE: SOUTH: Large-Scale Medical Informatics for Patient Care Coordination and Engagement
-
批准号:1636933
-
项目类别:Standard Grant
-
资助金额:$100.0万
-
财政年份:2016
-
负责人:Gari Clifford
-
依托单位:
Multi-scale markers of circadian rhythm changes for monitoring of mental health
-
批准号:EP/K020161/1
-
项目类别:Research Grant
-
资助金额:$11.34万
-
财政年份:2013
-
负责人:Gari Clifford
-
依托单位:
海外基金