Orchestrating privacy-protected big data analyses of data from different resources with R and DataSHIELD.

Orchestrating privacy-protected big data analyses of data from different resources with R and DataSHIELD.
复制标题

使用R和Datashield策划了来自不同资源的数据分析的大数据分析。

DOI:
10.1371/journal.pcbi.1008880
复制
发表时间:
2021-03
影响因子:
4.3
通讯作者:
González JR
González JR
中科院分区:
生物学2区
文献类型:
--
作者:
Marcon Y;Bishop T;Avraam D;Escriba-Montagut X;Ryser-Welch P;Wheater S;Burton P;González JR

文献摘要

参考文献

被引文献

相似文献

多个大型数据集的结合分析是健康和生物科学的共同目标。与分享结果的传统方法相关的分析性是,研究的数据在每个机构中都保留在服务器上,每个机构都可以控制他们的数据使用蛋白石,这是一个由流行病学研究使用的数据整合系统,由OBIBA开源项目在生物信息学的域中开发,但到目前为止,使用Datashield的大数据分析受OPAL中的存储格式限制,并且在DataShield restect中可用的dataShield Repantecter(我们的资源)允许使用新的建筑(“)。我们的新基础架构将帮助研究人员从现有数据中以受隐私保护的方式进行数据分析分享计划或项目,以帮助研究人员使用此框架,我们描述了选定的软件包(https://isglobal-brge.github.io/resource_bookdown)。 数据共享超出了任何一项研究的可能性,可以增加统计能力,并探索研究中的异质性。将数据界在同类财团中进行隐私保护的分析,对联合分析有真正的挑战。地理数据。我们还展示了GA4GH和EGA等基因组数据共享计划如何从我们的开发中受益。
Combined analysis of multiple, large datasets is a common objective in the health- and biosciences. Existing methods tend to require researchers to physically bring data together in one place or follow an analysis plan and share results. Developed over the last 10 years, the DataSHIELD platform is a collection of R packages that reduce the challenges of these methods. These include ethico-legal constraints which limit researchers’ ability to physically bring data together and the analytical inflexibility associated with conventional approaches to sharing results. The key feature of DataSHIELD is that data from research studies stay on a server at each of the institutions that are responsible for the data. Each institution has control over who can access their data. The platform allows an analyst to pass commands to each server and the analyst receives results that do not disclose the individual-level data of any study participants. DataSHIELD uses Opal which is a data integration system used by epidemiological studies and developed by the OBiBa open source project in the domain of bioinformatics. However, until now the analysis of big data with DataSHIELD has been limited by the storage formats available in Opal and the analysis capabilities available in the DataSHIELD R packages. We present a new architecture (“resources”) for DataSHIELD and Opal to allow large, complex datasets to be used at their original location, in their original format and with external computing facilities. We provide some real big data analysis examples in genomics and geospatial projects. For genomic data analyses, we also illustrate how to extend the resources concept to address specific big data infrastructures such as GA4GH or EGA, and make use of shell commands. Our new infrastructure will help researchers to perform data analyses in a privacy-protected way from existing data sharing initiatives or projects. To help researchers use this framework, we describe selected packages and present an online book (https://isglobal-brge.github.io/resource_bookdown). Data sharing enhances understanding of research results beyond what is possible from any single study. Data pooling across multiple studies increases statistical power and allows exploration of between-study heterogeneity. But, considerations related to ethico-legal and intellectual/commercial value regularly prevent or impede physical data sharing. DataSHIELD is designed to circumvent this problem. However, despite the growing confidence users have been placing in DataSHIELD to perform privacy-protected analyses of data in cohort consortia, there are real challenges to federated analytics. They include considering the wide range of data formats, and big data sources used, for example, in ‘omics-based research. This article describes the development and implementation of the new “resources” architecture in DataSHIELD that overcomes this limitation. We illustrate its value with real world examples related to genomics and geographical data. We also demonstrate how genomic data sharing initiatives such as GA4GH and EGA can benefit directly from our development. Our new infrastructure will help researchers to perform data analyses in a privacy-protected way from existing data sharing initiatives or projects.
DOI: 10.1093/bioinformatics/btz567
发表时间: 2019-12-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Gogarten, Stephanie M.;Sofer, Tamar;Conomos, Matthew P.
通讯作者: Conomos, Matthew P.
DOI: 10.1038/nmeth.3252
发表时间: 2015-02
期刊: Nature methods
影响因子: 48
作者:
Huber W;Carey VJ;Gentleman R;Anders S;Carlson M;Carvalho BS;Bravo HC;Davis S;Gatto L;Girke T;Gottardo R;Hahne F;Hansen KD;Irizarry RA;Lawrence M;Love MI;MacDonald J;Obenchain V;Oleś AK;Pagès H;Reyes A;Shannon P;Smyth GK;Tenenbaum D;Waldron L;Morgan M
通讯作者: Morgan M
DOI: 10.1136/bmj.g1464
发表时间: 2014-03-13
影响因子: 105.7
作者:
Burgoine, Thomas;Forouhi, Nita G.;Monsivais, Pablo
通讯作者: Monsivais, Pablo
DOI: 10.1038/s41588-020-0651-0
发表时间: 2020-07
期刊: Nature genetics
影响因子: 30.8
作者:
Bonomi L;Huang Y;Ohno-Machado L
通讯作者: Ohno-Machado L
DOI: 10.1093/bioinformatics/bts606
发表时间: 2012-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Zheng, Xiuwen;Levine, David;Weir, Bruce S.
通讯作者: Weir, Bruce S.