Secure Outsourcing of Genotype Imputation for Privacy-aware Genomic Analysis (RO1HE21)
Secure Outsourcing of Genotype Imputation for Privacy-aware Genomic Analysis (RO1HE21)
批准号:
10587347
负责人:
Arif Ozgun Harmanci
金额:
$58.32万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-05-17 至 2027-02-28
关键词:
All of Us Research ProgramAuthorization documentationAwarenessBenchmarkingCaliforniaCloud ServiceCollaborationsCommunicationCommunitiesComputing MethodologiesCuriositiesCustomDataData AnalysesData SecurityDatabasesDevelopmentDiscriminationEnsureEthicsFamilyFloridaGenealogyGeneticGenomeGenomicsGenotypeHospitalsHybridsIncentivesIndigenousIndividualLawsLearningLinkage DisequilibriumMapsMeasuresMethodsMichiganModelingNative-BornNatureOutsourced ServicesOutsourcingParticipantPatientsPatternPerformancePersonal Genetic InformationPersonsPhasePoliciesPopulationPopulation GeneticsPrivacyProcessPublic HealthPublishingQuality ControlRecreationReproducibilityResearchResearch PersonnelRestRiskRisk AssessmentSample SizeSecureSecurityServicesStatutes and LawsStigmatizationTechniquesTextTimeTrainingTrans-Omics for Precision MedicineTribesUnderrepresented PopulationsVariantauthoritycausal variantcloud basedcloud platformcomputing resourcesconvolutional neural networkcostdata sharingdesigndisease diagnosisencryptionflexibilitygenetic discriminationgenetic informationgenetic testinggenome sequencinggenome wide association studygenome-widegenomic datahigh dimensionalityimprovedin silicoinnovationmethod developmentnovelpopulation basedprivacy preservationprivacy protectionpublic health relevancetargeted sequencing
中文摘要
项目总结/摘要
人口规模的基因组测序项目,如1000个基因组,TOPMed和所有美国计划
将产生数百万个体的基因型数据。这一数字大幅增加,如果休闲
对来自家谱公司(如23andme)的遗传数据的使用进行了说明。共享和分析
这些数据对参与者的隐私造成了巨大的挑战。最近,黑客开始瞄准
家谱数据库,例如2020年GEDmatch的黑客攻击。由于规模大、维度高,
基因组数据,分析工作流程需要大量的计算资源。这激励了公司,医院,
和研究实验室使用第三方的外包服务来分析和解释基因组数据,
基因组数据存储在不可信的第三方服务器上。
在这个建议中,我们专注于安全外包的基因型填补,这是一个计算密集型
和大规模基因型分析的中心任务。基因型插补是缺失或低质量的预测
变异基因型,其使用一小组变异基因型,所述变异基因型使用例如基因分型
阵列、低覆盖率或靶向测序。这是分析原始基因组数据进行质量控制的重要步骤,
预测缺失的基因型、变体定相和关联的精细作图以鉴定因果变体。当
与稀疏阵列相结合,填补可以大大降低人口规模和家庭为基础的成本
基因分型例如,All of Us项目将依赖于定制的基因分型阵列Infinium Global Diversity
小组,以降低成本的基因分型数以百万计的个人。插补方法至关重要
来完成这项任务为了完成这些庞大的任务,插补方法需要大量的计算资源
并且经常被外包给第三方“估算服务器”。这些服务器将很快处理数以千计,如果没有
数以百万计的基因组,并存储敏感的基因组数据。不幸的是,这些服务并不严格安全
既不是来自未经授权的黑客,也不是来自有权访问服务器的好奇用户。有
迫切需要能够部署在甚至不受信任的第三方服务上的隐私感知估算方法
例如高性能云平台,以便在人口规模上安全地执行外包。
我们提出的方法使用最先进的同态加密,提供完美的基因组数据安全性
在运输、休息时,甚至在进行插补时。我们设计了新的高效的“加密-
保护研究参与者及其家人的“可靠的”方法和框架,
人口小组,即,代表性不足的人群。我们的基准测试表明,安全方法可以实现
即使在商用硬件上也具有高插补精度,其时间与最先进的非安全
方法.所提出的方法可以提供实用的人群规模的基因组隐私和安全的填补
协会研究。
英文摘要
Project Summary/Abstract
Population scale genome sequencing projects such as The 1000 Genomes, TOPMed, and All of US Program
will generate genotype data for millions of individuals. This number increases substantially if the recreational
usage of genetic data from genealogy companies, such as 23andme, is accounted for. Sharing and analyzing
this data create monumental challenges for the privacy of participants. Recently the hackers began targeting
genealogy databases such as the hacking of GEDmatch in 2020. Due to the large scale and high dimensions of
genomic data, analysis workflows require large computational resources. This incentivizes companies, hospitals,
and research labs to use outsourcing services from third parties to analyze and interpret genomic data such that
the genomic data is stored on untrusted 3rd party servers.
In this proposal, we focus on the secure outsourcing of genotype imputation, which is a computationally intensive
and central task in large-scale genotype analysis. Genotype imputation is the prediction of missing or low-quality
variant genotypes using a small set of variant genotypes that are measured using, for example, genotyping
arrays, low-coverage, or targeted sequencing. It is a vital step for analyzing raw genomic data for quality control,
predicting missing genotypes, variant phasing, and fine mapping of associations to identify causal variants. When
combined with sparse arrays, imputation can greatly reduce the cost of population-scale and family-based
genotyping. For example, the All of Us Project will rely on a custom genotyping array, Infinium Global Diversity
Panel, to decrease the cost of genotyping millions of individuals. Imputation methods will be of vital importance
for this task. To perform these enormous tasks, the imputation methods require large computational resources
and are often outsourced to 3rd party “imputation servers”. These servers will soon process thousands, If not
millions, of genomes and store sensitive genomic data. Unfortunately, these services are not strictly secure
neither from unauthorized hackers nor from curious users who have authorized access to the servers. There is
an urgent need for privacy-aware imputation methods that can be deployed on even untrusted 3rd party services
such as high-performance cloud platforms so that outsourcing can be safely performed at population scale.
Our proposed methods use state-of-the-art homomorphic encryption that provides perfect genomic data security
while in transit, at rest, and even while imputation is being performed. We design new and efficient “encryption-
amenable” methods and frameworks for protecting the study participants and their families, and for protecting
the population panels, i.e., underrepresented populations. Our benchmarks show that secure methods achieve
high imputation accuracy even on commodity hardware with comparable time as the state-of-the-art non-secure
methods. Proposed methods can provide practical population-scale genomic privacy and security for imputation
and association studies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金