Attacks on genetic privacy via uploads to genealogical databases

Attacks on genetic privacy via uploads to genealogical databases
复制标题

DOI:
10.7554/elife.51810
复制
发表时间:
2020-01-07
期刊:
影响因子:
7.7
通讯作者:
Coop, Graham
Coop, Graham
中科院分区:
生物学1区
文献类型:
--
作者:
Edge, Michael D.;Coop, Graham

文献摘要

被引文献

相似文献

直接面向消费者 (DTC) 的遗传学服务越来越受欢迎,拥有数千万客户。一些 DTC 家谱服务允许用户上传基因数据来搜索亲属,这些亲属被确定为具有相同州 (IBS) 区域基因组的人。在这里,我们描述了攻击者可以通过上传多个数据集来学习数据库基因型的方法。例如,上传大约 900 个基因组的对手可以在欧洲血统中位数人基因组的 82% 的 SNP 位点上恢复至少一个等位基因。在使用非定相基因型检测 IBS 片段的数据库中,大约 100 个伪造的上传可以揭示足够的遗传信息,以便进行全基因组遗传估算。我们在 GEDmatch 数据库中提供了概念验证演示,并提出了防止我们描述的漏洞的对策。
Direct-to-consumer (DTC) genetics services are increasingly popular, with tens of millions of customers. Several DTC genealogy services allow users to upload genetic data to search for relatives, identified as people with genomes that share identical by state (IBS) regions. Here, we describe methods by which an adversary can learn database genotypes by uploading multiple datasets. For example, an adversary who uploads approximately 900 genomes could recover at least one allele at SNP sites across up to 82% of the genome of a median person of European ancestries. In databases that detect IBS segments using unphased genotypes, approximately 100 falsified uploads can reveal enough genetic information to allow genome-wide genetic imputation. We provide a proof-of-concept demonstration in the GEDmatch database, and we suggest countermeasures that will prevent the exploits we describe.