Assessing transcriptomic reidentification risks using discriminative sequence models.

Assessing transcriptomic reidentification risks using discriminative sequence models.
复制标题

DOI:
10.1101/gr.277699.123
复制
发表时间:
2023-07
期刊:
影响因子:
7
通讯作者:
Cho, Hyunghoon
Cho, Hyunghoon
中科院分区:
生物学1区
文献类型:
--
作者:
Sadhuka, Shuvom;Fridman, Daniel;Berger, Bonnie;Cho, Hyunghoon

文献摘要

相似文献

基因表达数据提供了对遗传变异的功能影响的分子见解,例如,通过表达数量性状基因座(eQTL)。随着对基因型和基因表达之间关联的理解不断加深,人们越来越担心基因表达谱可能与另一个数据集中相同个体的基因型谱相匹配,这被称为链接攻击。先前的研究表明,由于模型假设的限制,这种风险只能分析一小部分独立的eQTL,从而无法完全理解这种风险的全部程度。为了解决这一挑战,我们引入了判别序列模型(DSM),一种新的概率框架,用于预测基于基因表达数据的基因型序列。通过对基因组区域中所有已知eQTL的联合分布进行建模,DSM提高了链接攻击的能力,并对连锁不平衡和冗余预测信号进行了必要的校准。与现有方法相比,DSM在一系列攻击场景和数据集(包括多达22,288人)中的链接准确性更高,这表明DSM有助于发现以前研究忽略的大量额外风险。我们的工作提供了一个统一的框架,用于评估共享转录组学之外的各种组学数据集的隐私风险。
Gene expression data provide molecular insights into the functional impact of genetic variation, for example, through expression quantitative trait loci (eQTLs). With an improving understanding of the association between genotypes and gene expression comes a greater concern that gene expression profiles could be matched to genotype profiles of the same individuals in another data set, known as a linking attack. Prior works show such a risk could analyze only a fraction of eQTLs that is independent owing to restrictive model assumptions, leaving the full extent of this risk incompletely understood. To address this challenge, we introduce the discriminative sequence model (DSM), a novel probabilistic framework for predicting a sequence of genotypes based on gene expression data. By modeling the joint distribution over all known eQTLs in a genomic region, DSM improves the power of linking attacks with necessary calibration for linkage disequilibrium and redundant predictive signals. We show greater linking accuracy of DSM compared with existing approaches across a range of attack scenarios and data sets including up to 22,288 individuals, suggesting that DSM helps uncover a substantial additional risk overlooked by previous studies. Our work provides a unified framework for assessing the privacy risks of sharing diverse omics data sets beyond transcriptomics.