SYNTHEMA: Synthetic generation of hematological data over federated computing frameworks
SYNTHEMA: Synthetic generation of hematological data over federated computing frameworks
批准号:
10054278
负责人:
金额:
$52.42万
依托单位国家:
英国
项目类别:
EU-Funded
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
血液病是由血细胞、淋巴器官和凝血因子的定量或定性异常引起的一大组疾病。尽管其中大多数(约74%)是罕见的,但全球受HD影响的患者总数很重要,给医疗保健系统和社会带来了相当大的经济负担。尽管在国家和欧盟层面存在多个合作研究小组,但由于每种疾病的患者数量相对较少,以及大量不相关的临床实体,目前的临床方法通常无效,特别是对于最罕见的疾病。SYNTHEMA旨在建立一个跨境数据中心,开发和验证基于人工智能的创新技术,用于临床数据匿名化和合成数据生成(SDG),以解决数据稀缺和碎片化问题,并扩大RHD中符合GDPR的研究基础。该项目将重点关注两个代表性的RHD用例:镰状细胞病(SCD)和急性髓性白血病(AML)。SYNTHEMA将开发一个联邦学习(FL)基础设施,配备安全多方计算(SMPC)和差分隐私(DF)协议,连接临床中心,带来标准化,可互操作的多模式数据集和来自学术界和中小企业的计算中心。该框架将用于训练开发的算法,并以隐私保护的方式执行基于SMPC的全局模型聚合。将验证所得数据的临床价值、统计效用和剩余隐私风险。该项目将制定法律的和道德框架,以保证在收集和处理与健康有关的个人数据时的隐私设计,并实现道德明智的算法共同创造。项目成果,包括管道,标准和数据,将向医疗保健,学术界和行业领域的利益相关者公开提供,并为现有的罕见疾病登记做出贡献。
英文摘要
Haematological diseases (HDs) are a large group of disorders resulting from quantitative or qualitative abnormalities of blood cells, lymphoid organs and coagulation factors. Despite most of them (~74%) are rare, the overall number of HD affected patients worldwide is important, placing a considerable economic burden on healthcare systems and societies. Despite the existence of several collaborative research groups at national and EU level, current clinical approaches are often ineffective, particularly for rarest conditions, due to the relatively low number of patients per disease and the high number of unconnected clinical entities. SYNTHEMA aims to establish a cross-border data hub where to develop and validate innovative AI-based techniques for clinical data anonymisation and synthetic data generation (SDG), to tackle the scarcity and fragmentation of data and widen the basis for GDPR-compliant research in RHDs. The project will focus on two representative RHD use cases: sickle-cell disease (SCD) and acute myeloid leukaemia (AML). SYNTHEMA will develop a federated learning (FL) infrastructure, equipped with secure multiparty computation (SMPC) and differential privacy (DF) protocols, connecting clinical centres bringing standardised, interoperable multimodal datasets and computing centres from academia and SME. This framework will be utilised to train the developed algorithms and perform SMPC-based global model aggregation in a privacy-preserving fashion. The resulting data will be validated for their clinical value, statistical utility and residual privacy risks. The project will develop legal and ethical frameworks to guarantee privacy by-design in the collection and processing of health-related personal data and attain an ethics-wise algorithm co-creation. Project outcomes, including pipelines, standards and data, will be made openly available to stakeholders in the healthcare, academia and industry field, and contribute to existing rare disease registries.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金