Novel methods to improve the utility of genomics summary statistics
Novel methods to improve the utility of genomics summary statistics
批准号:
10646125
负责人:
Nathan L Tintle
金额:
$41.22万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-19 至 2025-08-31
关键词:
AccelerationAddressAgeAgingAlzheimer&aposs DiseaseClinicalCognitiveCohort StudiesComputer softwareComputing MethodologiesConfidentiality of Patient InformationDataData ScienceData SetDatabasesDevelopmentDiseaseDocumentationEducational workshopElectronic Health RecordEnsureEtiologyFatty AcidsFundingFutureGeneticGenetic MarkersGenetic VariationGenetic studyGenomeGenomicsGenotypeHealthHeartHumanImpaired cognitionIndividualLinear ModelsMediatingMediationMeta-AnalysisMethodologyMethodsModelingNational Human Genome Research InstituteOutcomeOutcomes ResearchPhenotypePolyunsaturated Fatty AcidsPrivacyPythonsResearchResearch PersonnelRoleStatistical MethodsTestingTrainingValidationWorkbiobankclinically relevantcohortcostdata accessdata privacydietarydigital repositoriesflexibilitygenetic epidemiologygenetic variantgenome sciencesgenomic datahuman diseasehuman genome sequencingimprovedinnovationinterestmembernovelopen sourceopen source toolresponsesexstatisticssurvival outcometoolworking group
中文摘要
自测序以来的二十年里,反复的实验和计算突破
人类基因组为理解人类疾病的病因学提供了前所未有的机会。
基因组学数据成本的降低意味着研究人员现在有可能获得完整的基因组
数十万人的序列信息,通过大型数据库广泛访问这些数据
电子健康记录(EHR)和生物库的存储库。然而,相互关联的计算、统计
关于如何利用这些数据来研究基因变异的贡献,隐私问题仍然存在
到常见的疾病。重要的是,有必要进行方法创新,以最大限度地减少计算量
复杂性和尊重数据隐私问题,同时最大限度地提高数据访问和效用。站在…的最前列
这些创新是利用不可单独识别的汇总统计数据的计算方法
根据BIOBANK/EHR数据进行计算,以最大限度地了解下游功能和临床应用。一
新出现的一组汇总统计数据是来自独立回归模型的点估计和标准误差
表型(Y)对个体基因型(Xi)的影响,有时协变量调整有限(例如,年龄、性别
和主成分(PC))。任何一组预计算汇总统计数据的关键限制不是
能够预测此类统计数据的所有可能的下游用途。例如,研究人员可能希望使用:
(A)与预计算分析中考虑的协变量不同的协变量集合,(B)遗传变异集合
预测因子,而不是单一标记和(C)作为现有功能的替代表型定义
变量(例如,研究人员想知道具有临床重要性的表型𝑌𝑌𝐶𝐶,但只有Pre-
计算𝑌𝑌1、𝑌𝑌2、…的汇总统计信息,𝑌𝑌𝑘𝑘,其中𝑌𝑌𝐶𝐶=𝑓𝑓(𝑌𝑌1,𝑌𝑌2,…,𝑌𝑌𝑘𝑘)。在这个项目中,我们将(1)开发一个
计算高效的框架,用于评估具有临床相关表型的遗传变异
汇总统计数据并应用这些方法对临床相关表型进行协调分析
在多队列研究中使用汇总统计和(2)验证这些方法的可行性
目前正在探索基因变异对认知结果的作用的两个相关财团的创新
饮食中多不饱和脂肪酸(PUFA)水平的潜在调节和/或调节作用。初步
方法将在开放源码工具(R/python包)中实现,还将涉及广泛的测试
基于广泛的临床相关表型的模拟和真实数据。工作将为舞台搭建舞台
对于未来的R01项目,以提供额外的方法扩展、更广泛的测试和
全面传播。
英文摘要
The repeated experimental and computational breakthroughs in the two decades since the sequencing of the
human genome have provided an unprecedented opportunity to understand the etiology of human diseases.
The diminishing cost of genomics data means it is now possible for researchers to obtain complete genome
sequence information on hundreds of thousands of individuals, with widespread access to those data via large
repositories of electronic health records (EHRs) and biobanks. However, interrelated computational, statistical
and privacy questions remain for about how to leverage these data to study the contribution of genetic variation
to common diseases. Importantly, there is a need for methodological innovation to minimize computational
complexity and respect data privacy concerns while maximizing data access and utility. At the forefront of
these innovations are computational methods that leverage non-individually identifiable summary statistics pre-
computed on biobank/EHR data to maximize downstream functional understanding and clinical utility. One
emerging set of summary statistics are point estimates and standard errors from separate regression models
of a phenotype (Y) on individual genotypes (Xi), sometimes with limited covariate adjustment (e.g., Age, Sex
and principal components (PCs)). A key limitation of any set of pre-computed summary statistics is not being
able to anticipate all possible downstream uses of such statistics. For example, researchers may want to use:
(a) different sets of covariates than those considered in pre-computed analyses, (b) sets of genetic variants as
predictors, instead of single markers and (c) alternative phenotype definitions that are functions of existing
variables (e.g., a researcher want to know about a phenotype, 𝑌𝑌𝐶𝐶, of clinical importance, but only has pre-
computed summary statistics on 𝑌𝑌1, 𝑌𝑌2, … , 𝑌𝑌𝑘𝑘, where 𝑌𝑌𝐶𝐶 = 𝑓𝑓( 𝑌𝑌1, 𝑌𝑌2, … , 𝑌𝑌𝑘𝑘)). In this project we will (1) develop a
computationally efficient framework to evaluate genetic variants with clinically relevant phenotypes using
summary statistics and apply these methods to perform harmonized analyses of clinically relevant phenotypes
in multi-cohort studies using summary statistics and (2) validate the feasibility of these methodological
innovations within two related consortia currently exploring the role genetic variants on cognitive outcomes and
the potential moderating and/or mediating role of dietary polyunsaturated fatty acid (PUFA) levels. Preliminary
methods will be implemented in open-source tools (R/python packages), and will also involve extensive testing
on both simulated and real data across a wide range of clinically relevant phenotypes. Work will set the stage
for a future R01 project to provide additional methodological expansion, more widespread testing and
comprehensive dissemination.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Wastewater data integration and modelling to accurately predict community and organizational outbreaks due to viral pathogens
-
批准号:10481536
-
项目类别:
-
资助金额:$25.96万
-
财政年份:2022
-
负责人:Nathan L Tintle
-
依托单位:
Wastewater data integration and modelling to accurately predict community and organizational outbreaks due to viral pathogens
-
批准号:10768053
-
项目类别:
-
资助金额:$5.5万
-
财政年份:2022
-
负责人:Nathan L Tintle
-
依托单位:
Large-scale data integration and harmonization to accurately predict sites facing future health-based drinking water crises
-
批准号:10253600
-
项目类别:
-
资助金额:$25.66万
-
财政年份:2021
-
负责人:Nathan L Tintle
-
依托单位:
Analyzing the behavior and interpreting the results of gene based tests of rare variant association
-
批准号:9099474
-
项目类别:
-
资助金额:$38.59万
-
财政年份:2012
-
负责人:Nathan L Tintle
-
依托单位:
Analyzing the behavior and interpreting the results of gene based tests of rare v
-
批准号:8367623
-
项目类别:
-
资助金额:$39.16万
-
财政年份:2012
-
负责人:Nathan L Tintle
-
依托单位:
Analyzing the behavior and interpreting the results of gene based tests of rare variant association
-
批准号:9813293
-
项目类别:
-
资助金额:$35.27万
-
财政年份:2012
-
负责人:Nathan L Tintle
-
依托单位:
Evaluating the Cost Effectiveness of Alternative Sample Designs for Genetic Assoc
-
批准号:7841342
-
项目类别:
-
资助金额:$1.93万
-
财政年份:2009
-
负责人:Nathan L Tintle
-
依托单位:
Evaluating the Cost Effectiveness of Alternative Sample Designs for Genetic Assoc
-
批准号:8264409
-
项目类别:
-
资助金额:$7.48万
-
财政年份:2008
-
负责人:Nathan L Tintle
-
依托单位:
Evaluating the Cost Effectiveness of Alternative Sample Designs for Genetic Assoc
-
批准号:7363067
-
项目类别:
-
资助金额:$19.35万
-
财政年份:2008
-
负责人:Nathan L Tintle
-
依托单位:
海外基金