Infrastructure and distributed learning methodology for privacy-preserving multi-centric rapid learning health care: euroCAT.

Infrastructure and distributed learning methodology for privacy-preserving multi-centric rapid learning health care: euroCAT.
复制标题

DOI:
10.1016/j.ctro.2016.12.004
复制
发表时间:
2017-06
影响因子:
3.1
通讯作者:
Lambin P
Lambin P
中科院分区:
医学3区
文献类型:
--
作者:
Deist TM;Jochems A;van Soest J;Nalbantov G;Oberije C;Walsh S;Eble M;Bulens P;Coucke P;Dries W;Dekker A;Lambin P

文献摘要

被引文献

相似文献

在3个国家和地区的5家放射诊所开发和实施了IT基础设施。大数据基础设施和分布式学习研究的原则证明。对分布式数据执行学习算法的通用框架。个性化医疗的机器学习应用高度依赖于对足够数据的访问。对于个性化的放射肿瘤学,需要获取代表整个癌症患者群体的变异的数据集,并使用该数据集来学习预测模型。确保数据隐私的道德和法律界限阻碍了研究机构之间的合作。我们假设,在没有可识别的患者数据离开放射诊所的情况下,数据共享是可能的,并且在分布式数据集上构建机器学习应用是可行的。我们在三个国家(比利时、德国和荷兰)的五个放射诊所开发并实施了IT基础设施。我们在这里为未来的大数据基础设施和分布式学习研究提供了一个原则证明。所有五个地点的肺癌患者数据都被收集并存储在当地数据库中。使用乘数交替方向法(ADMM)从分布式数据库学习示例性支持向量机(SVM)模型,以预测放射治疗后呼吸困难分级。通过五次交叉验证(在四个站点学习,在第五个站点验证)中的曲线下面积(AUC)来评估区分性能。将分布式学习算法的性能与集中式学习进行了比较,集中式学习是对所有研究所的数据集进行联合分析的。Eurocat基础设施已在三个国家的五个放射诊所成功实施。支持向量机模型可以从分布在所有五家诊所的数据中学习。此外,该基础设施还提供了对分布式数据执行学习算法的通用框架。Eurocat网络的不断扩大将促进放射肿瘤学中的机器学习。由此获得的具有足够差异的更大数据集将为可推广的预测模型和个性化医学铺平道路。
Developed and implemented IT infrastructure in 5 radiation clinics across 3 countries. Proof-of-principle for ‘big data’ infrastructure and distributed learning studies. General framework to execute learning algorithms on distributed data. Machine learning applications for personalized medicine are highly dependent on access to sufficient data. For personalized radiation oncology, datasets representing the variation in the entire cancer patient population need to be acquired and used to learn prediction models. Ethical and legal boundaries to ensure data privacy hamper collaboration between research institutes. We hypothesize that data sharing is possible without identifiable patient data leaving the radiation clinics and that building machine learning applications on distributed datasets is feasible. We developed and implemented an IT infrastructure in five radiation clinics across three countries (Belgium, Germany, and The Netherlands). We present here a proof-of-principle for future ‘big data’ infrastructures and distributed learning studies. Lung cancer patient data was collected in all five locations and stored in local databases. Exemplary support vector machine (SVM) models were learned using the Alternating Direction Method of Multipliers (ADMM) from the distributed databases to predict post-radiotherapy dyspnea grade . The discriminative performance was assessed by the area under the curve (AUC) in a five-fold cross-validation (learning on four sites and validating on the fifth). The performance of the distributed learning algorithm was compared to centralized learning where datasets of all institutes are jointly analyzed. The euroCAT infrastructure has been successfully implemented in five radiation clinics across three countries. SVM models can be learned on data distributed over all five clinics. Furthermore, the infrastructure provides a general framework to execute learning algorithms on distributed data. The ongoing expansion of the euroCAT network will facilitate machine learning in radiation oncology. The resulting access to larger datasets with sufficient variation will pave the way for generalizable prediction models and personalized medicine.