Privacy-Preserving Data Sharing for Genome-Wide Association Studies

Privacy-Preserving Data Sharing for Genome-Wide Association Studies
复制标题

DOI:
10.29012/jpc.v5i1.629
复制
发表时间:
2012-05
期刊:
The Journal of privacy and confidentiality
影响因子:
--
通讯作者:
Caroline Uhler;A. Slavkovic;S. Fienberg
Caroline Uhler;A. Slavkovic;S. Fienberg
中科院分区:
其他
文献类型:
--
作者:
Caroline Uhler;A. Slavkovic;S. Fienberg

文献摘要

相似文献

全基因组关联研究(GWAS)专注于发现与重大疾病等性状相关的遗传变异,通常通过测量单核苷酸多态性(SNP)之间的关联,即,单核苷酸的DNA序列变异和特定疾病。一项典型的研究比较了患有疾病的个体(病例)和没有疾病的类似个体(对照组)的DNA。对于特定性状,此类研究的输出通常由最显著SNP的χ2统计量或p值组成,包括其次要等位基因频率(即,在病例和对照中观察到的最低等位基因频率)。在一篇震惊遗传学界的文章中,Homer等人[13]声称,在某些条件下,他们可以使用统计方法来“准确和鲁棒地[解决]”一个具有已知基因型的个体在混合DNA样本中的存在,其中只有次要等位基因频率(MAF)是已知的。他们的方法比较了特定个体的MAF与参考人群中MAF的分布以及测试人群中MAF的分布;然后他们使用t检验来评估该个体是否是测试人群的一部分。虽然被专门提出用于法医环境,只是次要的用于破坏隐私,但Homer等人的“攻击”似乎普遍适用。作为参考人群,可以使用来自HapMap项目1的公开可用的SNP数据,其由来自大小从45到90个个体不等的四个人群的SNP数据组成。请注意,HapMap数据集不包含任何关于个体健康状况的信息。对于测试人群,可以使用GWAS中的病例,其中包含基因型数据和疾病状态。在文章出现之前[13],GWAS中病例的平均MAF和对照组的平均MAF通常是公开的。作为对Homer等人[13]的回应,Braun等人[3]表明,他们提出的测试在很大程度上依赖于这样的假设,即测试人群、参考人群和所考虑的特定人的基因型是来自相同基础人群的样本,并且研究中使用的SNP是独立的(即,不存在连锁不平衡)。这些假设在实践中通常不符合,因此,Homer等人的攻击导致高误报率,参见例如,Braun et al. [3]. Homer等人的其他批评建议了识别问题的替代公式,声称可以加强攻击,或者建议不同的方法来保护数据,例如,见[6,14,15,16,18,19,21,23,27]。尽管Homer等人[13]对GWAS参与者隐私的攻击存在明显的局限性,并且我们认为他们的统计声明存在争议,并且夸大了其统计性质,但NIH立即从开放获取数据库中删除了所有汇总结果,例如病例和对照组的平均MAF值,卡方(χ2)统计量和p值(参见Couzin [7]以及Zerhouni和Nabel [26])。美国国立卫生研究院的政策至今仍然有效。2每一位想要访问这些数据集的研究人员都需要经过一个复杂的批准程序。对于在GWAS研究中没有可靠记录的计算机科学家、数学家或统计学家来说,这是一个特别困难的障碍。在这里,我们提出了允许在不损害个人隐私的情况下发布聚合GWAS数据的方法,并且在许多方面完全避开了Homer等人[13]和其他人关于GWAS数据库脆弱性的主张的有效性的辩论。我们的GWAS隐私保证利用了最近由加密社区引入的差分隐私概念(例如,Dwork等人[10])。差分隐私提供了一个严格的隐私定义,在任意外部信息存在的情况下,有意义的隐私保证。我们的贡献如下:我们提出了一种方法,平均MAF的情况下,并在GWAS的控制,而不损害个人的隐私释放。我们计算e-differentially private χ2-statistics和p-values,并提供了一个差异私有算法,用于发布最相关的SNP的这些统计数据。癌症、心脏病和糖尿病等疾病是由各种基因和可能的环境相互作用引起的。检测与特定表型相关的SNP之间的这种相互作用(即,上位性)是GWAS的主要目标。大多数用于发现上位性的方法基于两阶段方法:(1)过滤所有SNP,例如,使用χ2统计或简单逻辑回归,以将潜在相互作用的SNP减少到少量;(2)进一步检查达到相互作用的一些阈值的基因座。例如,Park和Hastie [20]使用一种惩罚逻辑回归的形式来检测少数SNP上的基因-基因相互作用。通过将[1]和[5]的工作适应于这种方法,我们导出了GWAS的隐私保护方法,其中两阶段方法中的两个阶段都满足e-差分隐私。第二节介绍了基本问题和相关定义。在第3节中,我们介绍了用于发布e-差异私有MAF、χ2统计量和p值的方法,在第4节中,我们基于模拟研究和涉及685只犬的犬毛长度的GWAS研究的数据评估了它们的统计效用。在第5节中,我们提出了一种基于惩罚方法的logistic回归的差异私有方法来寻找全基因组关联。
Genome-wide association studies (GWAS) focus on finding genetic variations associated with traits such as major diseases often by measuring associations between single-nucleotide polymorphisms (SNPs), i.e., DNA sequence variations at single nucleotides, and a particular disease. A typical study compares the DNA of individuals with the disease (cases) and similar individuals without (controls). For a specific trait, the output of such studies often consists of the χ2-statistics or the p-values for the most significant SNPs including their minor allele frequencies (i.e., the lowest allele frequency observed for the cases and the controls). In an article that shocked the genetics community, Homer et al. [13] claimed that, under certain conditions, they could use statistical methods to “accurately and robustly [resolve]” the presence of an individual with known genotype in a mix of DNA samples from which only the minor allele frequencies (MAFs) are known. Their approach compared the MAFs of a specific individual to the distribution of MAFs in a reference population and the distribution of MAFs in a test population; they then used a t-test to assess if the individual was part of the test population. Although proposed specifically for use in a forensic context and only secondarily for breaking privacy, the Homer et al. [13] “attack” appeared to be generally applicable. As a reference population one might use the publicly available SNP data from the HapMap project1 which consists of SNP data from four populations varying in size from 45 to 90 individuals. Note that the HapMap data set does not contain any information regarding the health status of the individuals. For the test population one might use the cases in GWAS, which contain both genotype data and disease status. Before the appearance of the article [13], the averaged MAFs of the cases and the averaged MAFs of the controls in a GWAS were typically publicly available. In response to Homer et al. [13], Braun et al. [3] showed that their proposed test depends heavily on the assumption that the genotypes of the test population, the reference population, and the specific person under consideration are samples from the same underlying population, and that the SNPs used in the study are independent (i.e., that there is no linkage disequilibrium present). These assumptions are usually not met in practice, and as a consequence, the Homer et al. [13] attack lead to a high false-positive rate, see e.g., Braun et al. [3]. Other critiques of Homer et al. suggested alternative formulations of the identification problem, claimed to strengthen the attack, or suggested different ways to protect the data, e.g., see [6, 14, 15, 16, 18, 19, 21, 23, 27]. Despite the apparent limitations of the Homer et al. [13] attack on the privacy of GWAS participants and the controversial and, we believe, exaggerated nature of their statistical claims, NIH immediately removed from open-access databases all aggregate results such as values of averaged MAFs over cases and controls, chi-square (χ2)-statistics, and p-values (see Couzin [7] and Zerhouni and Nabel [26]). The NIH policy remains in effect today.2 Every researcher who wants to gain access to any of these data sets needs to go through an elaborate approval process. This is a particularly difficult obstacle for computer scientists, mathematicians, or statisticians who do not have a credible record in GWAS research. Here we propose methods which allow for the release of aggregate GWAS data without compromising an individual’s privacy, and in many ways totally bystep the debate on the validity of the claims by Homer et al. [13] and others on the vulnerability of GWAS databases. Our GWAS privacy guarantees utilize the concept of differential privacy, recently introduced by the cryptographic community (e.g., Dwork et al. [10]). Differential privacy provides a rigorous definition of privacy with meaningful privacy guarantees in the presence of arbitrary external information. Our contributions are as follows: We propose a method for the release of the averaged MAFs for the cases and for the controls in GWAS without compromising an individual’s privacy. We compute e-differentially private χ2-statistics and p-values and provide a differentially private algorithm for releasing these statistics for the most relevant SNPs. Conditions such as cancer, heart disease, and diabetes are caused by the interaction of various genes and possibly the environment. Detecting such interaction among SNPs related to a specific phenotype (i.e., epistasis) is a main goal of GWAS. Most methods for finding epistasis are based on a two-stage approach: (1) Filtering all SNPs, e.g., using χ2-statistics or a simple logistic regression, to reduce the potentially interacting SNPs to a small number; (2) Further examining the loci achieving some threshold for interactions. For example, Park and Hastie [20] use a form of penalized logistic regression to test for detecting gene-gene interactions on a small number of SNPs. By adapting the work of [1] and [5] to this methodology, we derive a privacy-preserving method for GWAS, where both stages in the two-stage approach satisfy e-differential privacy. Section 2 describes the basic problem and relevant definitions. In Section 3, we present methods for releasing e-differentially private MAFs, χ2-statistics, and p-values, and in Section 4 we evaluate their statistical utility on data based on a simulation study and on a GWAS study of canine hair length involving 685 dogs. In Section 5, we propose a differentially-private method for finding genome-wide associations based on a penalized approach to logistic regression.