Privacy-Preserving Data Sharing for Genome-Wide Association Studies
Privacy-Preserving Data Sharing for Genome-Wide Association Studies
复制标题
DOI:
10.29012/jpc.v5i1.629
复制
发表时间:
2012-05
期刊:
影响因子:
--
通讯作者:
Caroline Uhler;A. Slavkovic;S. Fienberg
中科院分区:
文献类型:
--
作者:
Caroline Uhler;A. Slavkovic;S. Fienberg
Genome-wide association studies (GWAS) focus on finding genetic variations associated with traits such as major diseases often by measuring associations between single-nucleotide polymorphisms (SNPs), i.e., DNA sequence variations at single nucleotides, and a particular disease. A typical study compares the DNA of individuals with the disease (cases) and similar individuals without (controls). For a specific trait, the output of such studies often consists of the χ2-statistics or the p-values for the most significant SNPs including their minor allele frequencies (i.e., the lowest allele frequency observed for the cases and the controls). In an article that shocked the genetics community, Homer et al. [13] claimed that, under certain conditions, they could use statistical methods to “accurately and robustly [resolve]” the presence of an individual with known genotype in a mix of DNA samples from which only the minor allele frequencies (MAFs) are known. Their approach compared the MAFs of a specific individual to the distribution of MAFs in a reference population and the distribution of MAFs in a test population; they then used a t-test to assess if the individual was part of the test population. Although proposed specifically for use in a forensic context and only secondarily for breaking privacy, the Homer et al. [13] “attack” appeared to be generally applicable. As a reference population one might use the publicly available SNP data from the HapMap project1 which consists of SNP data from four populations varying in size from 45 to 90 individuals. Note that the HapMap data set does not contain any information regarding the health status of the individuals. For the test population one might use the cases in GWAS, which contain both genotype data and disease status. Before the appearance of the article [13], the averaged MAFs of the cases and the averaged MAFs of the controls in a GWAS were typically publicly available. In response to Homer et al. [13], Braun et al. [3] showed that their proposed test depends heavily on the assumption that the genotypes of the test population, the reference population, and the specific person under consideration are samples from the same underlying population, and that the SNPs used in the study are independent (i.e., that there is no linkage disequilibrium present). These assumptions are usually not met in practice, and as a consequence, the Homer et al. [13] attack lead to a high false-positive rate, see e.g., Braun et al. [3]. Other critiques of Homer et al. suggested alternative formulations of the identification problem, claimed to strengthen the attack, or suggested different ways to protect the data, e.g., see [6, 14, 15, 16, 18, 19, 21, 23, 27]. Despite the apparent limitations of the Homer et al. [13] attack on the privacy of GWAS participants and the controversial and, we believe, exaggerated nature of their statistical claims, NIH immediately removed from open-access databases all aggregate results such as values of averaged MAFs over cases and controls, chi-square (χ2)-statistics, and p-values (see Couzin [7] and Zerhouni and Nabel [26]). The NIH policy remains in effect today.2 Every researcher who wants to gain access to any of these data sets needs to go through an elaborate approval process. This is a particularly difficult obstacle for computer scientists, mathematicians, or statisticians who do not have a credible record in GWAS research. Here we propose methods which allow for the release of aggregate GWAS data without compromising an individual’s privacy, and in many ways totally bystep the debate on the validity of the claims by Homer et al. [13] and others on the vulnerability of GWAS databases. Our GWAS privacy guarantees utilize the concept of differential privacy, recently introduced by the cryptographic community (e.g., Dwork et al. [10]). Differential privacy provides a rigorous definition of privacy with meaningful privacy guarantees in the presence of arbitrary external information. Our contributions are as follows: We propose a method for the release of the averaged MAFs for the cases and for the controls in GWAS without compromising an individual’s privacy. We compute e-differentially private χ2-statistics and p-values and provide a differentially private algorithm for releasing these statistics for the most relevant SNPs. Conditions such as cancer, heart disease, and diabetes are caused by the interaction of various genes and possibly the environment. Detecting such interaction among SNPs related to a specific phenotype (i.e., epistasis) is a main goal of GWAS. Most methods for finding epistasis are based on a two-stage approach: (1) Filtering all SNPs, e.g., using χ2-statistics or a simple logistic regression, to reduce the potentially interacting SNPs to a small number; (2) Further examining the loci achieving some threshold for interactions. For example, Park and Hastie [20] use a form of penalized logistic regression to test for detecting gene-gene interactions on a small number of SNPs. By adapting the work of [1] and [5] to this methodology, we derive a privacy-preserving method for GWAS, where both stages in the two-stage approach satisfy e-differential privacy. Section 2 describes the basic problem and relevant definitions. In Section 3, we present methods for releasing e-differentially private MAFs, χ2-statistics, and p-values, and in Section 4 we evaluate their statistical utility on data based on a simulation study and on a GWAS study of canine hair length involving 685 dogs. In Section 5, we propose a differentially-private method for finding genome-wide associations based on a penalized approach to logistic regression.