Identity inference of genomic data using long-range familial searches

Identity inference of genomic data using long-range familial searches
复制标题

DOI:
10.1126/science.aau4832
复制
发表时间:
2018-11-09
期刊:
影响因子:
56.9
通讯作者:
Carmi, Shai
Carmi, Shai
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Erlich, Yaniv;Shor, Tal;Carmi, Shai

文献摘要

被引文献

相似文献

消费者基因组学数据库已经达到数百万人的规模。最近,执法当局利用其中一些数据库,通过远房亲戚来查明嫌疑人。使用消费者基因组学测试的128万个体的基因组数据,我们研究了这项技术的力量。我们预计,大约60%的欧洲血统的个人搜索将导致第三堂兄弟或更接近的匹配,这在理论上允许他们使用人口统计标识符识别。此外,在不久的将来,这项技术几乎可以涉及任何欧洲血统的美国人。我们证明,该技术也可以识别公共测序项目的研究参与者。在这些结果的基础上,我们提出了一个潜在的缓解策略和政策含义,为人类受试者的研究。
Consumer genomics databases have reached the scale of millions of individuals. Recently, law enforcement authorities have exploited some of these databases to identify suspects via distant familial relatives. Using genomic data of 1.28 million individuals tested with consumer genomics, we investigated the power of this technique. We project that about 60% of the searches for individuals of European descent will result in a third-cousin or closer match, which theoretically allows their identification using demographic identifiers. Moreover, the technique could implicate nearly any U.S. individual of European descent in the near future. We demonstrate that the technique can also identify research participants of a public sequencing project. On the basis of these results, we propose a potential mitigation strategy and policy implications for human subject research.