Caught you: threats to confidentiality due to the public release of large-scale genetic data sets.

Caught you: threats to confidentiality due to the public release of large-scale genetic data sets.
复制标题

DOI:
10.1186/1472-6939-11-21
复制
发表时间:
2010-12-29
期刊:
影响因子:
2.7
通讯作者:
Wjst M
Wjst M
中科院分区:
人文科学2区
文献类型:
--
作者:
Wjst M

文献摘要

参考文献

被引文献

相似文献

大规模的基因数据集经常与其他研究小组分享,甚至在互联网上发布,以便进行二次分析。研究参与者通常不会被告知此类数据共享,因为数据集在剥离个人标识符后被认为是匿名的。然而,遗传数据集匿名的假设是脆弱的,因为遗传数据本质上是自我识别的。两种类型的重新识别是可能的:“Netflix”类型和“配置文件”类型。“Netflix”类型需要另一个小的遗传数据集,通常少于100个SNP,但包括个人标识符。第二个数据集可能来自另一个临床检查,剩余样本的研究或法医测试。当合并到主要未识别集时,它将重新识别该个体的所有样本。即使手头没有第二组数据,也可以制定一种“分析”战略,从样本收集中提取尽可能多的信息。从沿着的身体特征和疾病的预测的种族亚群的识别,哮喘儿童的情况下,作为一个现实生活中的例子来说明这种方法。根据补充信息的程度,很有可能至少可以从匿名数据集中识别出几个人。然而,任何重新鉴定都可能对研究参与者造成潜在伤害,因为它将向公众释放个体遗传疾病风险。
Large-scale genetic data sets are frequently shared with other research groups and even released on the Internet to allow for secondary analysis. Study participants are usually not informed about such data sharing because data sets are assumed to be anonymous after stripping off personal identifiers. The assumption of anonymity of genetic data sets, however, is tenuous because genetic data are intrinsically self-identifying. Two types of re-identification are possible: the "Netflix" type and the "profiling" type. The "Netflix" type needs another small genetic data set, usually with less than 100 SNPs but including a personal identifier. This second data set might originate from another clinical examination, a study of leftover samples or forensic testing. When merged to the primary, unidentified set it will re-identify all samples of that individual. Even with no second data set at hand, a "profiling" strategy can be developed to extract as much information as possible from a sample collection. Starting with the identification of ethnic subgroups along with predictions of body characteristics and diseases, the asthma kids case as a real-life example is used to illustrate that approach. Depending on the degree of supplemental information, there is a good chance that at least a few individuals can be identified from an anonymized data set. Any re-identification, however, may potentially harm study participants because it will release individual genetic disease risks to the public.
DOI: 10.1038/sj.ejhg.5201701
发表时间: 2006-11-01
影响因子: 5.2
作者:
Shepperd, Sasha;Farndon, Peter;Rose, Peter
通讯作者: Rose, Peter
DOI: 10.1038/nature06014
发表时间: 2007-07-26
期刊: NATURE
影响因子: 64.8
作者:
Moffatt, Miriam F.;Kabesch, Michael;Cookson, William O. C.
通讯作者: Cookson, William O. C.
DOI: 10.1136/jamia.2009.000026
发表时间: 2010-03-01
影响因子: 6.4
作者:
Benitez, Kathleen;Malin, Bradley
通讯作者: Malin, Bradley
DOI: 10.1371/journal.pgen.1000167
发表时间: 2008-08-29
期刊: PLoS genetics
影响因子: 4.5
作者:
Homer N;Szelinger S;Redman M;Duggan D;Tembe W;Muehling J;Pearson JV;Stephan DA;Nelson SF;Craig DW
通讯作者: Craig DW
DOI: 10.1016/j.ajhg.2009.01.018
发表时间: 2009-02-13
影响因子: 9.8
作者:
Gitschier, Jane
通讯作者: Gitschier, Jane