Quantification of private information leakage from phenotype-genotype data: linking attacks.

Quantification of private information leakage from phenotype-genotype data: linking attacks.
复制标题

DOI:
10.1038/nmeth.3746
复制
发表时间:
2016-03
期刊:
影响因子:
48
通讯作者:
Gerstein M
Gerstein M
中科院分区:
生物学1区
文献类型:
--
作者:
Harmanci A;Gerstein M

文献摘要

被引文献

相似文献

基因组隐私的研究传统上集中在使用DNA变异识别个人。相比之下,分子表型数据,如基因表达水平,通常被认为没有这样的识别信息。虽然它们没有明确的基因型信息,但对手可以使用公开的基因型-表型相关性(例如,表达数量性状基因座(eQTL))将表型与基因型统计学联系起来。当使用高维数据(许多表达水平)时,这种链接可以是准确的,并且由此产生的链接可以揭示敏感信息,例如,患有癌症的个体。在这里,我们开发的框架,量化泄漏的个人特征信息表型数据集。这些可以用于在发布之前估计来自大型数据集的泄漏。我们还提出了一个一般的三步程序,实际上实例化连接攻击和特定的攻击使用离群基因表达水平,这是简单而准确的。最后,我们描述了这种离群值攻击在不同情况下的有效性。
Studies on genomic privacy have traditionally focused on identifying individuals using DNA variants. In contrast, molecular phenotype data, such as gene expression levels, are generally assumed free of such identifying information. Although there is no explicit genotypic information in them, adversaries can statistically link phenotypes to genotypes using publicly available genotype-phenotype correlations, for instance, expression quantitative trait loci (eQTLs). This linking can be accurate when high-dimensional data (many expression levels) are used, and the resulting links can then reveal sensitive information, for example, an individual having cancer. Here, we develop frameworks for quantifying the leakage of individual characterizing information from phenotype datasets. These can be used for estimating the leakage from large datasets before release. We also present a general three-step procedure for practically instantiating linking attacks and a specific attack using outlier gene-expression levels that is simple yet accurate. Finally, we describe the effectiveness of this outlier attack under different scenarios.