Identifying Personal Genomes by Surname Inference

Identifying Personal Genomes by Surname Inference
复制标题

DOI:
10.1126/science.1229566
复制
发表时间:
2013-01-18
期刊:
影响因子:
56.9
通讯作者:
Erlich, Yaniv
Erlich, Yaniv
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Gymrek, Melissa;McGuire, Amy L.;Erlich, Yaniv

文献摘要

被引文献

相似文献

共享没有标识符的测序数据集已成为基因组学中的常见做法。在这里,我们报告说,姓氏可以从个人基因组中恢复的Y染色体上的短串联重复序列(Y-STR)和查询娱乐遗传家谱数据库。我们表明,姓氏与其他类型的元数据,如年龄和状态的组合,可以用来三角定位的目标的身份。这种技术的一个关键特征是它完全依赖于免费的、可公开访问的互联网资源。我们定量分析了美国男性的识别概率。我们进一步证明了这种技术的可行性,通过追溯与高概率的多个参与者在公共测序项目的身份。
Sharing sequencing data sets without identifiers has become a common practice in genomics. Here, we report that surnames can be recovered from personal genomes by profiling short tandem repeats on the Y chromosome (Y-STRs) and querying recreational genetic genealogy databases. We show that a combination of a surname with other types of metadata, such as age and state, can be used to triangulate the identity of the target. A key feature of this technique is that it entirely relies on free, publicly accessible Internet resources. We quantitatively analyze the probability of identification for U.S. males. We further demonstrate the feasibility of this technique by tracing back with high probability the identities of multiple participants in public sequencing projects.