Synthetic data in medical research.

Synthetic data in medical research.
复制标题

DOI:
10.1136/bmjmed-2022-000167
复制
发表时间:
2022
期刊:
BMJ medicine
影响因子:
--
通讯作者:
Harron, Katie
Harron, Katie
中科院分区:
其他
文献类型:
--
作者:
Kokosi, Theodora;Harron, Katie

文献摘要

参考文献

被引文献

相似文献

医疗和保健研究对获取个人层面的高质量数据的需求正在不断增长。收集整个人群的电子健康记录数据可以帮助生成真实世界的证据,并可用于一系列次要目的,包括测试新假设以及开发和评估不同的方法和统计方法。对主要研究数据(例如来自临床试验的数据)的二次分析也很有价值,例如对个体参与者数据进行荟萃分析。然而,一些复杂的隐私要求使得访问这些数据具有挑战性。 2 电子健康记录或临床试验数据中包含的信息高度敏感,访问这些数据集可能是一个昂贵且漫长的过程。 3 数据隐私和保护法规是获取这些数据以进行医疗保健和医学研究的主要障碍。 4 匿名化(删除潜在可识别变量)是提供数据的一种方法;然而,密集的匿名化可能会降低数据的质量,使其不再适合目的。 5 例如,向数据添加随机噪声会降低精度并导致置信区间变大。对匿名数据进行的几次重新识别尝试都取得了成功,但损害了公众和监管机构对此类方法的信任。 6 7 例如,一项研究表明,可以通过匹配公开的患者级别数据中的信息、归因于从报纸获得的信息以及直接联系这些患者来识别患者。 6使用来自大量人群的临床试验和电子健康记录的信息有可能使医疗和保健研究受益,并使得寻求新的数据访问方法势在必行。一种解决方案是使用所谓的合成数据或人工数据,它们提供原始数据的真实表示
Demand to access high quality data at the individual level for medical and healthcare research is growing. Electronic health record data collected on whole populations can help to generate real world evidence and can be used for a range of secondary purposes, including testing new hypotheses and developing and evaluating different methodological and statistical approaches. Secondary analysis of primary research data, such as from clinical trials, 1 is also valuable—for example, to conduct meta-analyses of individual participant data. However, several complex privacy requirements make accessing these data challenging. 2 Information contained in electronic health records or in clinical trial data are highly sensitive and access to these datasets can be an expensive and lengthy process. 3 Data privacy and protection regulations are the main barriers to accessing these data for healthcare and medical research. 4 Anonymisation (where potentially identifiable variables are removed) is one way to make data available; however, intensive anonymisation can degrade the data to the extent that it is no longer fit for purpose. 5 For example, adding random noise to the data reduces precision and leads to larger confidence intervals. Several reidentification attempts on anonymised data have been successful and have harmed public and regulators’ trust in such methods. 6 7 For instance, one study showed that patients could be identified by matching information from patient level data that was publicly available, attributing information obtained from newspapers, and contacting those patients directly. 6Use of information from clinical trials and electronic health records of large populations has the potential to benefit medical and healthcare research and makes seeking new approaches to data access imperative. One solution is to use so-called synthetic data, or artificial data, which provide a realistic representation of the original
DOI: 10.1136/bmjopen-2020-043497
发表时间: 2021-04-16
期刊: BMJ open
影响因子: 2.9
作者:
Azizi Z;Zheng C;Mosquera L;Pilote L;El Emam K;GOING-FWD Collaborators
通讯作者: GOING-FWD Collaborators
DOI: 10.1038/s41746-020-00353-9
发表时间: 2020-11-09
影响因子: 15.2
作者:
Tucker A;Wang Z;Rotalinti Y;Myles P
通讯作者: Myles P
DOI: 10.1109/jbhi.2020.2980262
发表时间: 2020-08-01
影响因子: 7.7
作者:
Yoon, Jinsung;Drumright, Lydia N.;van der Schaar, Mihaela
通讯作者: van der Schaar, Mihaela
DOI: 10.14778/3231751.3231757
发表时间: 2018-06-01
影响因子: 2.5
作者:
Park, Noseong;Mohammadi, Mahmoud;Kim, Youngmin
通讯作者: Kim, Youngmin
DOI: 10.1111/rssa.12358
发表时间: 2018-06-01
影响因子: 2
作者:
Snoke, Joshua;Raab, Gillian M.;Slavkovic, Aleksandra
通讯作者: Slavkovic, Aleksandra