Data harmonization and federated analysis of population-based studies: the BioSHaRE project.

Data harmonization and federated analysis of population-based studies: the BioSHaRE project.
复制标题

DOI:
10.1186/1742-7622-10-12
复制
发表时间:
2013-11-21
影响因子:
2.3
通讯作者:
Fortier I
Fortier I
中科院分区:
其他
文献类型:
--
作者:
Doiron D;Burton P;Marcon Y;Gaye A;Wolffenbuttel BHR;Perola M;Stolk RP;Foco L;Minelli C;Waldenberger M;Holle R;Kvaløy K;Hillege HL;Tassé AM;Ferretti V;Fortier I

文献摘要

被引文献

相似文献

在国际研究项目中跨研究中心进行大规模人口研究的个人层面数据汇集面临许多障碍。BioSHaRE(欧盟卓越研究生物库标准化和协调)项目旨在通过建立一个研究人员合作小组和开发数据协调、数据库集成和联邦数据分析工具来解决这些问题。BioSHaRE项目招募了6个欧洲国家的8项基于人群的研究。通过讲习班、电话会议和电子通信,参与的调查人员确定了一套96个变量,以协调一致,以回答感兴趣的研究问题。利用每个研究的问卷、标准操作程序和数据字典,评估了协调的潜力。只要认为协调是可能的,就在开源软件基础设施中开发和实施处理算法,以将特定于研究的数据转换为目标(即协调)格式。位于欧洲各研究中心服务器上的统一数据集通过联邦数据库系统相互连接,进行统计分析。回顾性协调导致73%的匹配产生共同的格式变量(8项研究中的96个目标变量)。经过认证的调查人员现在可以对存储在分布式服务器上的协调数据集进行复杂的统计分析,而无需使用DataSHIELD方法实际共享个人层面的数据。新的基于因特网的网络技术和数据库管理系统提供了一种有效和安全的方式来支持协作、多中心研究。这个试点项目的结果表明,鉴于参与研究之间强有力的合作关系,在允许每个研究保留对个人层面数据的完全控制的同时,无缝地共同分析国际协调的研究数据库是可能的。我们鼓励流行病学、公共卫生和社会科学领域的其他合作研究网络利用本文提供的开源工具。
Individual-level data pooling of large population-based studies across research centres in international research projects faces many hurdles. The BioSHaRE (Biobank Standardisation and Harmonisation for Research Excellence in the European Union) project aims to address these issues by building a collaborative group of investigators and developing tools for data harmonization, database integration and federated data analyses. Eight population-based studies in six European countries were recruited to participate in the BioSHaRE project. Through workshops, teleconferences and electronic communications, participating investigators identified a set of 96 variables targeted for harmonization to answer research questions of interest. Using each study’s questionnaires, standard operating procedures, and data dictionaries, harmonization potential was assessed. Whenever harmonization was deemed possible, processing algorithms were developed and implemented in an open-source software infrastructure to transform study-specific data into the target (i.e. harmonized) format. Harmonized datasets located on server in each research centres across Europe were interconnected through a federated database system to perform statistical analysis. Retrospective harmonization led to the generation of common format variables for 73% of matches considered (96 targeted variables across 8 studies). Authenticated investigators can now perform complex statistical analyses of harmonized datasets stored on distributed servers without actually sharing individual-level data using the DataSHIELD method. New Internet-based networking technologies and database management systems are providing the means to support collaborative, multi-center research in an efficient and secure manner. The results from this pilot project show that, given a strong collaborative relationship between participating studies, it is possible to seamlessly co-analyse internationally harmonized research databases while allowing each study to retain full control over individual-level data. We encourage additional collaborative research networks in epidemiology, public health, and the social sciences to make use of the open source tools presented herein.