Creating longitudinal datasets for linked administrative data research using synthetic data
Creating longitudinal datasets for linked administrative data research using synthetic data
批准号:
ES/V005448/1
负责人:
Katie Harron
金额:
$20.56万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
已结题
起止时间:
2021 至 --
中文摘要
行政数据在为公共政策提供信息方面具有巨大潜力。然而,由于数据访问、链接和隐私保护方面的限制,这一潜力尚未实现。治理程序和审批导致数据访问的时间跨度很长,限制很严,这可能会危及公共资助的研究。一种解决方案是生成合成数据,保留原始来源的统计属性,但不对应于任何真实的个人,也不会带来隐私风险。这些数据可以广泛共享,使研究人员能够了解数据结构、开发分析计划和算法并测试不同的模型。这可以与申请访问链接的行政数据集同时进行,从而简化研究过程。最后的提炼和分析将在真实的数据上进行。我们的研究将测试创建合成链接的行政数据集的方法的可行性。我们将比较两种现有的方法:“Synthpop”,用于创建苏格兰纵向研究的合成版本,“Simulacrum”,用于创建国家癌症登记处的合成版本,新方法“Jomo”,基于最近的方法学发展,用于填补缺失数据。我们将评估这些方法使用的范例连接第三次全国性态度和生活方式调查(Natsal-3)的两个行政数据集:医院事件统计(HES)和全国学生数据库(NPD)。Natsal-3是世界上最大的性人口为基础的行为调查之一,收集了15000名参与者在2010 - 2012年期间的数据。HES包含英格兰所有NHS医院的就诊信息,可以对程序和诊断进行详细分析。NPD包含有关在英格兰公立学校就读的学生的信息,包括学业成绩,缺勤和特殊教育需求。Natsal-3、HES和NPD之间的联系将提供一个独特的机会,以更深入地了解性健康和生殖健康的社会、行为和生物学方面,并为实施性健康干预措施提供证据。(因为它们都具有不同的结构和特征),这取决于这些方法生成的数据对原始数据的表现。我们还将申请链接数据的批准,以i)探索在合成复杂的链接数据时是否需要任何额外的考虑因素,以及ii)生成可以与研究人员更广泛共享的链接数据的合成版本。合成数据的质量和可用性在很大程度上取决于数据生成模型和分析目的。然而,确定所有相关变量以及这些变量之间可能的依赖关系或相互作用是一项高度资源密集型工作。因此,合成数据生成的挑战之一是了解是否存在合成数据的通用版本可能足以用于某些目的的情况,或者是否总是需要定制的合成数据集(针对特定的研究问题)。我们将通过与数据提供者和研究人员接触,并确定产生可接受的输出所需的两者之间沟通的性质和实用性,来探索这种平衡。我们亦会与公众接触,征询他们对使用合成数据的意见。基于一组示例研究问题,我们将生成合成数据,并比较不同方法的可行性和输出。为了评估合成数据在多大程度上代表了真实的数据,我们将比较合成数据与真实的数据的特征和统计推断。根据我们的研究结果,我们将制定关于适当使用合成数据的指导方针。
英文摘要
Administrative data hold great potential for informing public policy. However, this potential is not yet being realised due to restrictions around data access, linkage, and privacy protection. Governance procedures and approvals lead to long timescales and tight restrictions on data access, which can jeopardise publicly funded research.One solution is to generate synthetic data that preserve the statistical properties of the original sources, but do not correspond to any real individuals or pose privacy risks. These data could be widely shared, allowing researchers to understand the data structures, develop analysis plans and algorithms, and test out different models. This could be done in parallel to applying for access to linked administrative datasets, streamlining the research process. Final refinements and analyses would be conducted on the real data.Our study will test the feasibility of approaches for creating synthetic linked administrative datasets. We will compare two existing methods: 'Synthpop', used to create synthetic versions of the Scottish Longitudinal Study, and 'Simulacrum', used to create synthetic versions of the National Cancer Registry, with a new approach 'Jomo', based on recent methodological developments for the imputation of missing data. We will evaluate these approaches using an exemplar of linking the third National Survey of Sexual Attitudes and Lifestyles (Natsal-3) to two administrative datasets: Hospital Episode Statistics (HES) and the National Pupil Database (NPD).Natsal-3 is one of the largest sexual population-based behaviour surveys in the world and collected data from 15000 participants during 2010-2012. HES contains information on attendances to all NHS hospitals in England, allowing detailed analysis of procedures and diagnoses. NPD contains information on pupils attending state schools in England, including school achievement, absences, and special educational needs. Linkage between Natsal-3, HES and NPD will provide a unique opportunity to gain a deeper understanding of the social, behavioural and biological aspects of sexual and reproductive health, and to generate evidence to inform implementation of sexual health interventions.We will first compare different methods for generating synthetic versions of the three datasets separately (since all have different structures and characteristics), based on how well the data generated by these methods represent the original data. We will also apply for approvals to link the data, to i) explore whether there are any additional considerations needed when synthesising complex, linked data, and ii) generate synthetic versions of the linked data that can be shared with researchers more widely.The quality and usability of synthetic data is highly dependent on the data generation model and the purpose of analysis. However, identifying all relevant variables and possible dependencies or interactions between these is highly resource intensive. One of the challenges for synthetic data generation is therefore understanding whether there are situations in which generic versions of synthetic data may be sufficient for some purposes, or whether bespoke synthetic datasets (tailored to a specific research problem) are always required. We will explore this balance by engaging with data providers and researchers and determining the nature and practicality of communication between the two that is required to produce acceptable outputs. We will also engage with the public to seek their views on the use of synthetic data. Based on a set of exemplar research questions, we will generate synthetic data and compare feasibility and outputs from different approaches. To evaluate how well the synthetic data represent the real data, we will compare characteristics and statistical inferences from the synthetic data with those from the real data. Based on our findings, we will generate guidelines on the appropriate use of synthetic data.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1136/bmjmed-2022-000167
发表时间:
2022
期刊:
BMJ medicine
影响因子:
--
作者:
[Kokosi, Theodora, Harron, Katie]
通讯作者:
Harron, Katie
Enhancement of ECHILD with a mother-child and Unique Property Reference number link
-
批准号:ES/X000427/1
-
项目类别:Research Grant
-
资助金额:$34.64万
-
财政年份:2022
-
负责人:Katie Harron
-
依托单位:
Linking health and education data for research to improve outcomes for children in England - Supplement
-
批准号:ES/X003663/1
-
项目类别:Research Grant
-
资助金额:$49.43万
-
财政年份:2022
-
负责人:Katie Harron
-
依托单位:
国内基金
海外基金
精神分裂症进程中非对称性活跃脑结构改变的磁共振研究
-
批准号:81171275
-
项目类别:面上项目
-
资助金额:14.0万元
-
批准年份:2011
-
负责人:邓伟
-
依托单位: