MetaSeq: Privacy Preserving Meta-Analysis of Sequencing-Based Association Studies

MetaSeq: Privacy Preserving Meta-Analysis of Sequencing-Based Association Studies
复制标题

DOI:
10.7916/d8fn1h8s
复制
发表时间:
2012-11
影响因子:
--
通讯作者:
Angad P. Singh;Samreen Zafer;I. Pe’er
Angad P. Singh;Samreen Zafer;I. Pe’er
中科院分区:
--
文献类型:
--
作者:
Angad P. Singh;Samreen Zafer;I. Pe’er

文献摘要

相似文献

人类遗传学最近从GWAS过渡到基于NGS数据的研究。对于GWAS来说,小的影响决定了大的样本量,通常通过跨联盟交换汇总统计数据的荟萃分析来实现。NGS研究GroupWise检验,以确定每个基因上多个潜在因果等位基因的关联性。它们受到类似的权力限制,因此可能也会求助于荟萃分析。当在数据交换过程中考虑遗传信息的隐私时,这个问题就出现了。许多NGS关联评分方案依赖于每个变体的频率,因此需要交换已测序变体的身份。因为这种变种通常很少见,可能会暴露其运营商的身份,并危及隐私。因此,我们开发了MetaSeq协议,这是一种由多个协作方对全基因组测序数据进行荟萃分析的协议,为所有各方共享的每个基因的罕见变异打分。我们解决了统计稀有、已测序等位基因的频率计数的挑战,用于对测序数据进行荟萃分析,而不披露等位基因身份和计数,从而保护样本身份。这种明显自相矛盾的信息交换是通过加密手段实现的。关键的想法是各方对基因和变种的身份进行加密。当他们传输有关病例和控制中的频率计数的信息时,交换的数据不会传达突变的身份,因此不会暴露携带者身份。该交换依赖于第三方,该第三方被信任遵循协议,但不被信任来了解原始数据。我们展示了这种方法对来自多个研究的公开可获得的外显子组测序数据的适用性,模拟了强大的荟萃分析的表型信息。MetaSeq软件以开放源码的形式公开提供。
Human genetics recently transitioned from GWAS to studies based on NGS data. For GWAS, small effects dictated large sample sizes, typically made possible through meta-analysis by exchanging summary statistics across consortia. NGS studies groupwise-test for association of multiple potentially-causal alleles along each gene. They are subject to similar power constraints and therefore likely to resort to meta-analysis as well. The problem arises when considering privacy of the genetic information during the data-exchange process. Many scoring schemes for NGS association rely on the frequency of each variant thus requiring the exchange of identity of the sequenced variant. As such variants are often rare, potentially revealing the identity of their carriers and jeopardizing privacy. We have thus developed MetaSeq, a protocol for meta-analysis of genome-wide sequencing data by multiple collaborating parties, scoring association for rare variants pooled per gene across all parties. We tackle the challenge of tallying frequency counts of rare, sequenced alleles, for metaanalysis of sequencing data without disclosing the allele identity and counts, thereby protecting sample identity. This apparent paradoxical exchange of information is achieved through cryptographic means. The key idea is that parties encrypt identity of genes and variants. When they transfer information about frequency counts in cases and controls, the exchanged data does not convey the identity of a mutation and therefore does not expose carrier identity. The exchange relies on a 3rd party, trusted to follow the protocol although not trusted to learn about the raw data. We show applicability of this method to publicly available exome-sequencing data from multiple studies, simulating phenotypic information for powerful meta-analysis. The MetaSeq software is publicly available as open source.