课题基金 / 基金详情

Advancing Federated Learning of Neural Networks for Medical Imaging

Advancing Federated Learning of Neural Networks for Medical Imaging
推进医学成像神经网络的联合学习
批准号:
2594573
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
机器学习(ML)算法,学习检测数据模式的计算方法,通过实现快速准确的医学图像分析,有望改善疾病的诊断和治疗。最先进的机器学习方法,深度神经网络(DNN),通常被训练为使用手动标记的数据来识别模式,例如成对的医学扫描和相应的手动生成的标签,这些标签描述了扫描显示的病理。这种标记的医学数据是有限的,因为临床医生的注释是昂贵的。此外,由于隐私问题,将来自世界各地临床中心的数据汇总到一个中央数据库中通常是不可行的。因此,训练数据库很小,无法捕获临床实践中的真实的异质性(罕见病变、不同扫描仪等)。因此,在如此有限的数据上训练的DNN不能很好地推广,这阻碍了它们在医疗保健中的采用。该项目将开发方法,使多个机构能够在其数据上进行协作和训练单个DNN,而无需将它们集中在一个计算节点中。这个框架被称为多个计算节点(机构)之间的DNN的联合学习(FL)。用FL训练的模型可以通过从世界各地收集的各种数据库中学习来更好地推广。这可以为改善疾病诊断和治疗带来强大而可靠的ML工具。FL有可能成为大规模国际ML医疗研究的标准范式。然而,有多种技术挑战阻碍了其有效使用。该项目解决以下问题:a)在不同临床中心采集的数据具有异质性,例如由于不同的患者人口统计数据或采集扫描仪。当在具有这种系统差异的数据库之间执行FL时,模型优化是次优的。这是因为常见的优化方法假设数据是相同和独立分布的(iid),这在FL设置中是不正确的。该项目将开发用于非iid数据的FL的优化算法以提高其有效性。B)当应用于呈现与用于训练的数据不同的特征的数据时,现有DNN的性能是不可靠的。我们将研究如何识别和建模用于FL的数据库之间的变化因素(例如来自不同机构),从而能够推断部署后的预期变化,以提高模型的泛化能力。c)标签通常在医疗保健中受到限制,而未标记的数据则非常丰富。FL方法主要是为使用标签进行学习而设计的。该项目将使用未标记的数据开发FL,使任何机构都能够在合作联盟中提供其未标记的数据,从而使模型能够更好地捕捉世界各地真实的数据异质性。这项研究是及时的,将推动医学图像分析和ML领域的发展。FL在医疗保健中的价值已经得到证实,引起了极大的兴趣,但技术挑战限制了其使用。从非iid和未标记的数据中学习是ML中长期存在的挑战。因此,本项目的结果是有价值的医学图像分析,但也感兴趣的其他领域。该项目福尔斯属于EPSRC医学成像研究领域和医疗保健技术主题。它的最终目标是创建有效的FL工具,使医学成像界能够进行协作研究,改善疾病诊断和治疗。这项研究是在牛津大学生物医学工程研究所与大数据研究所合作进行的。它是由现有的合作与帝国理工学院伦敦和剑桥大学促进,并将寻求建立新的在英国和国际。
英文摘要
Machine Learning (ML) algorithms, computational methods that learn to detect patterns in data, promise to improve diagnosis and treatment of disease by enabling fast and accurate medical image analysis. State of the art ML methods, Deep Neural Networks (DNNs), are commonly trained to identify patterns using manually labelled data, such as pairs of medical scans and corresponding manally-generated labels that describe what pathology the scans show. Such labelled medical data are limited because annotation by clinicians is expensive. Moreover, aggregating data from clinical centres across the world in one central database is often infeasible due to privacy concerns. As a result, training databases are small and do not capture the real heterogeneity in clinical practice (rare pathologies, different scanners, etc). Consequently, DNNs trained on such limited data do not generalize well, which hinders their adoption in healthcare. This project will develop methods that enable multiple institutions to collaborate and train a single DNN on their data, without the need to centrally aggregate them in one computational node. This framework is known as Federated Learning (FL) of DNNs between multiple computational nodes (institutions). Models trained with FL could potentially generalize better by learning from diverse databases collected across the world. This can lead to powerful and reliable ML tools for improved disease diagnosis and treatment.FL has the potential to become the standard paradigm for large-scale, international studies on ML for healthcare. There are multiple technical challenges, however, hindering its effective use. This project tackles the following:a) Data acquired at different clinical centres have heterogeneous characteristics, such as due to varying patient demographics or acquisition scanners. When FL is performed between databases with such systematic differences, model optimization is suboptimal. This is because common optimization methods assume the data are identically and independently distributed (iid), which is not true in an FL setting. This project will develop optimization algorithms for FL with non-iid data to improve its effectiveness.b) Performance of existing DNNs is unreliable when applied on data that present different characteristics from those used for training. We will investigate how to identify and model factors of variation between databases used for FL (e.g. from different institutions), enabling inference about expected variability after deployment, to improve model generalization.c) Labels are often limited in healthcare, whereas unlabelled data are abundant. FL methods have been primarily designed for learning using labels. This project will develop FL using unlabelled data, enabling any institution to provide their unlabelled data in a collaborative consortium, to allow models capture better the true data heterogeneity across the world.This research is timely and will advance medical image analysis and the field of ML. Value of FL in healthcare has been demonstrated previously, generating great interest, but technical challenges limit its use. Learning from non-iid and unlabelled data are long standing challenges in ML yet to be solved. Hence results by this project are valuable for medical image analysis but also of interest to other domains. This project falls within the EPSRC Medical Imaging research area and the Healthcare Technologies theme. Its ultimate goal is to create effective FL tools to enable the medical imaging community perform collaborative studies and improve disease diagnosis and treatment.This research is conducted at the University of Oxford within the Institute of Biomedical Engineering, in collaboration with the Big-Data Institute. It is facilitated by existing collaborations with Imperial College London and University of Cambridge, and will seek to establish new ones within UK and internationally.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金