Deterministic identification of specific individuals from GWAS results

Deterministic identification of specific individuals from GWAS results
复制标题

根据 GWAS 结果确定性识别特定个体

DOI:
10.1093/bioinformatics/btv018
复制
发表时间:
2015-06-01
期刊:
影响因子:
5.8
通讯作者:
Zhou, Shuigeng
Zhou, Shuigeng
中科院分区:
生物学3区
文献类型:
--
作者:
Cai, Ruichu;Hao, Zhifeng;Zhou, Shuigeng

文献摘要

被引文献

相似文献

动机:全基因组关联研究(GWASs)通常应用于人类基因组数据,以了解在统计上与某些疾病相关的因果基因组合。当研究公布了大量单核苷酸多态性的统计信息后,参与这些GWASs的患者可以被重新识别。然而,随后的研究发现,这种隐私攻击在理论上是可能的,但在现实环境中是不成功的,也不令人信服。结果:我们获得了第一个实用的隐私攻击,可以成功地从威康信托病例控制联盟(WTCCC)数据集中有限的已发布协会中识别特定的个人。对于超过25个随机选择的基因座计算的GWAS结果,我们的算法总是从WTCCC数据集中精确定位至少一个患者。此外,随着已发表的基因型数量的增加,重新鉴定的患者数量也在迅速增长。最后,我们讨论了阻止攻击的预防方法,从而提供了增强患者隐私的解决方案。
Motivation: Genome-wide association studies (GWASs) are commonly applied on human genomic data to understand the causal gene combinations statistically connected to certain diseases. Patients involved in these GWASs could be re-identified when the studies release statistical information on a large number of single-nucleotide polymorphisms. Subsequent work, however, found that such privacy attacks are theoretically possible but unsuccessful and unconvincing in real settings.Results: We derive the first practical privacy attack that can successfully identify specific individuals from limited published associations from the Wellcome Trust Case Control Consortium (WTCCC) dataset. For GWAS results computed over 25 randomly selected loci, our algorithm always pinpoints at least one patient from the WTCCC dataset. Moreover, the number of re-identified patients grows rapidly with the number of published genotypes. Finally, we discuss prevention methods to disable the attack, thus providing a solution for enhancing patient privacy.