Disclosure control of machine learning models from trusted research environments (TRE): New challenges and opportunities.

Disclosure control of machine learning models from trusted research environments (TRE): New challenges and opportunities.
复制标题

DOI:
10.1016/j.heliyon.2023.e15143
复制
发表时间:
2023-04
期刊:
影响因子:
4
通讯作者:
Jefferson, Emily
Jefferson, Emily
中科院分区:
综合性期刊4区
文献类型:
--
作者:
Mansouri-Benssassi, Esma;Rogers, Simon;Reel, Smarti;Malone, Maeve;Smith, Jim;Ritchie, Felix;Jefferson, Emily

文献摘要

参考文献

被引文献

相似文献

近年来,人工智能(AI)在医疗保健和医学领域的应用越来越多。为了能够访问个人数据,可信研究环境(TRE)(也称为安全港)提供了安全可靠的环境,研究人员可以在其中访问敏感的个人数据并开发AI(特别是机器学习(ML))模型。然而,目前很少有TREs支持ML模型的训练,部分原因是TREs在处理模型披露方面的实际决策指导存在差距。具体来说,ML模型的训练需要披露来自TREs的新类型的输出。尽管TREs有明确的统计输出披露政策,但训练模型一旦发布,个人训练数据的泄露程度尚不清楚。我们为普通观众回顾了不同类型的ML模型及其在医疗保健中的适用性。我们解释了训练ML模型的输出,以及经过训练的ML模型如何容易受到外部攻击,从而发现模型中编码的个人数据。我们提出了在从TREs训练和导出模型的背景下,对训练的ML模型进行披露控制的挑战。我们提供了可以在TREs中引入的见解和分析方法,以减轻在披露训练模型时隐私泄露的风险。虽然贸易效率报告中有具体的统计披露控制准则和政策,但这些准则和政策不能令人满意地解决这些新类型的产出要求,即,训练的ML模型新的跨学科研究机会在开发和调整政策和工具以安全地披露TREs的ML输出方面具有巨大的潜力。
Artificial intelligence (AI) applications in healthcare and medicine have increased in recent years. To enable access to personal data, Trusted Research Environments (TREs) (otherwise known as Safe Havens) provide safe and secure environments in which researchers can access sensitive personal data and develop AI (in particular machine learning (ML)) models. However, currently few TREs support the training of ML models in part due to a gap in the practical decision-making guidance for TREs in handling model disclosure. Specifically, the training of ML models creates a need to disclose new types of outputs from TREs. Although TREs have clear policies for the disclosure of statistical outputs, the extent to which trained models can leak personal training data once released is not well understood. We review, for a general audience, different types of ML models and their applicability within healthcare. We explain the outputs from training a ML model and how trained ML models can be vulnerable to external attacks to discover personal data encoded within the model. We present the challenges for disclosure control of trained ML models in the context of training and exporting models from TREs. We provide insights and analyse methods that could be introduced within TREs to mitigate the risk of privacy breaches when disclosing trained models. Although specific guidelines and policies exist for statistical disclosure controls in TREs, they do not satisfactorily address these new types of output requests; i.e., trained ML models. There is significant potential for new interdisciplinary research opportunities in developing and adapting policies and tools for safely disclosing ML outputs from TREs.
DOI: 10.2174/1573405616666200712180521
发表时间: 2021-01-01
影响因子: 1.4
作者:
Babu, K. Rajesh;Nagajaneyulu, P., V;Prasad, K. Satya
通讯作者: Prasad, K. Satya
DOI: 10.1109/msec.2021.3076443
发表时间: 2021-07-01
影响因子: 1.9
作者:
De Cristofaro, Emiliano
通讯作者: De Cristofaro, Emiliano
DOI: 10.1038/nature21056
发表时间: 2017-02-02
期刊: Nature
影响因子: 64.8
作者:
Esteva A;Kuprel B;Novoa RA;Ko J;Swetter SM;Blau HM;Thrun S
通讯作者: Thrun S
DOI: 10.1038/s41551-021-00751-8
发表时间: 2021-06
影响因子: 28.1
作者:
通讯作者: --
DOI: 10.1186/s12961-018-0308-y
发表时间: 2018-04-25
影响因子: 4
作者:
Chung Y;Salvador-Carulla L;Salinas-Pérez JA;Uriarte-Uriarte JJ;Iruin-Sanz A;García-Alonso CR
通讯作者: García-Alonso CR