SYNTHEMA: Synthetic generation of hematological data over federated computing frameworks
SYNTHEMA: Synthetic generation of hematological data over federated computing frameworks
批准号:
10054278
负责人:
金额:
$52.42万
依托单位国家:
英国
项目类别:
EU-Funded
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
血液病是由血细胞、淋巴器官和凝血因子的数量或质量异常引起的一大类疾病。尽管其中大多数(~74%)是罕见的,但全球受HD影响的患者总数是重要的,这给医疗体系和社会带来了相当大的经济负担。尽管在国家和欧盟层面上存在几个合作研究小组,但目前的临床方法往往无效,特别是对于最罕见的情况,因为每种疾病的患者数量相对较少,而临床实体的数量却很多。SYNTHEMA的目标是建立一个跨境数据中心,在那里开发和验证基于人工智能的创新技术,用于临床数据匿名化和合成数据生成(SDG),以解决数据稀缺和碎片化的问题,并扩大RHD中符合GDPR的研究的基础。该项目将侧重于两个具有代表性的风湿性心脏病使用案例:镰状细胞病(SCD)和急性髓系白血病(AML)。SYNTHEMA将开发一个联邦学习(FL)基础设施,配备安全多方计算(SMPC)和差异隐私(DF)协议,将临床中心连接起来,带来标准化的、可互操作的多模式数据集,以及来自学术界和中小企业的计算中心。这个框架将被用来训练开发的算法,并以保护隐私的方式执行基于SMPC的全球模型聚合。由此产生的数据将被验证其临床价值、统计效用和残留的隐私风险。该项目将制定法律和道德框架,以保障在收集和处理与健康有关的个人数据时的隐私权,并实现合乎道德的算法共同创造。项目成果,包括管道、标准和数据,将向医疗保健、学术界和行业领域的利益攸关方公开提供,并为现有的罕见疾病登记做出贡献。
英文摘要
Haematological diseases (HDs) are a large group of disorders resulting from quantitative or qualitative abnormalities of blood cells, lymphoid organs and coagulation factors. Despite most of them (~74%) are rare, the overall number of HD affected patients worldwide is important, placing a considerable economic burden on healthcare systems and societies. Despite the existence of several collaborative research groups at national and EU level, current clinical approaches are often ineffective, particularly for rarest conditions, due to the relatively low number of patients per disease and the high number of unconnected clinical entities. SYNTHEMA aims to establish a cross-border data hub where to develop and validate innovative AI-based techniques for clinical data anonymisation and synthetic data generation (SDG), to tackle the scarcity and fragmentation of data and widen the basis for GDPR-compliant research in RHDs. The project will focus on two representative RHD use cases: sickle-cell disease (SCD) and acute myeloid leukaemia (AML). SYNTHEMA will develop a federated learning (FL) infrastructure, equipped with secure multiparty computation (SMPC) and differential privacy (DF) protocols, connecting clinical centres bringing standardised, interoperable multimodal datasets and computing centres from academia and SME. This framework will be utilised to train the developed algorithms and perform SMPC-based global model aggregation in a privacy-preserving fashion. The resulting data will be validated for their clinical value, statistical utility and residual privacy risks. The project will develop legal and ethical frameworks to guarantee privacy by-design in the collection and processing of health-related personal data and attain an ethics-wise algorithm co-creation. Project outcomes, including pipelines, standards and data, will be made openly available to stakeholders in the healthcare, academia and industry field, and contribute to existing rare disease registries.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金