Novel methods to improve the utility of genomics summary statistics
Novel methods to improve the utility of genomics summary statistics
批准号:
10646125
负责人:
Nathan L Tintle
金额:
$41.22万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-19 至 2025-08-31
关键词:
AccelerationAddressAgeAgingAlzheimer&aposs DiseaseClinicalCognitiveCohort StudiesComputer softwareComputing MethodologiesConfidentiality of Patient InformationDataData ScienceData SetDatabasesDevelopmentDiseaseDocumentationEducational workshopElectronic Health RecordEnsureEtiologyFatty AcidsFundingFutureGeneticGenetic MarkersGenetic VariationGenetic studyGenomeGenomicsGenotypeHealthHeartHumanImpaired cognitionIndividualLinear ModelsMediatingMediationMeta-AnalysisMethodologyMethodsModelingNational Human Genome Research InstituteOutcomeOutcomes ResearchPhenotypePolyunsaturated Fatty AcidsPrivacyPythonsResearchResearch PersonnelRoleStatistical MethodsTestingTrainingValidationWorkbiobankclinically relevantcohortcostdata accessdata privacydietarydigital repositoriesflexibilitygenetic epidemiologygenetic variantgenome sciencesgenomic datahuman diseasehuman genome sequencingimprovedinnovationinterestmembernovelopen sourceopen source toolresponsesexstatisticssurvival outcometoolworking group
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The repeated experimental and computational breakthroughs in the two decades since the sequencing of the
human genome have provided an unprecedented opportunity to understand the etiology of human diseases.
The diminishing cost of genomics data means it is now possible for researchers to obtain complete genome
sequence information on hundreds of thousands of individuals, with widespread access to those data via large
repositories of electronic health records (EHRs) and biobanks. However, interrelated computational, statistical
and privacy questions remain for about how to leverage these data to study the contribution of genetic variation
to common diseases. Importantly, there is a need for methodological innovation to minimize computational
complexity and respect data privacy concerns while maximizing data access and utility. At the forefront of
these innovations are computational methods that leverage non-individually identifiable summary statistics pre-
computed on biobank/EHR data to maximize downstream functional understanding and clinical utility. One
emerging set of summary statistics are point estimates and standard errors from separate regression models
of a phenotype (Y) on individual genotypes (Xi), sometimes with limited covariate adjustment (e.g., Age, Sex
and principal components (PCs)). A key limitation of any set of pre-computed summary statistics is not being
able to anticipate all possible downstream uses of such statistics. For example, researchers may want to use:
(a) different sets of covariates than those considered in pre-computed analyses, (b) sets of genetic variants as
predictors, instead of single markers and (c) alternative phenotype definitions that are functions of existing
variables (e.g., a researcher want to know about a phenotype, 𝑌𝑌𝐶𝐶, of clinical importance, but only has pre-
computed summary statistics on 𝑌𝑌1, 𝑌𝑌2, … , 𝑌𝑌𝑘𝑘, where 𝑌𝑌𝐶𝐶 = 𝑓𝑓( 𝑌𝑌1, 𝑌𝑌2, … , 𝑌𝑌𝑘𝑘)). In this project we will (1) develop a
computationally efficient framework to evaluate genetic variants with clinically relevant phenotypes using
summary statistics and apply these methods to perform harmonized analyses of clinically relevant phenotypes
in multi-cohort studies using summary statistics and (2) validate the feasibility of these methodological
innovations within two related consortia currently exploring the role genetic variants on cognitive outcomes and
the potential moderating and/or mediating role of dietary polyunsaturated fatty acid (PUFA) levels. Preliminary
methods will be implemented in open-source tools (R/python packages), and will also involve extensive testing
on both simulated and real data across a wide range of clinically relevant phenotypes. Work will set the stage
for a future R01 project to provide additional methodological expansion, more widespread testing and
comprehensive dissemination.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Wastewater data integration and modelling to accurately predict community and organizational outbreaks due to viral pathogens
-
批准号:10481536
-
项目类别:
-
资助金额:$25.96万
-
财政年份:2022
-
负责人:Nathan L Tintle
-
依托单位:
Wastewater data integration and modelling to accurately predict community and organizational outbreaks due to viral pathogens
-
批准号:10768053
-
项目类别:
-
资助金额:$5.5万
-
财政年份:2022
-
负责人:Nathan L Tintle
-
依托单位:
Large-scale data integration and harmonization to accurately predict sites facing future health-based drinking water crises
-
批准号:10253600
-
项目类别:
-
资助金额:$25.66万
-
财政年份:2021
-
负责人:Nathan L Tintle
-
依托单位:
Analyzing the behavior and interpreting the results of gene based tests of rare variant association
-
批准号:9099474
-
项目类别:
-
资助金额:$38.59万
-
财政年份:2012
-
负责人:Nathan L Tintle
-
依托单位:
Analyzing the behavior and interpreting the results of gene based tests of rare v
-
批准号:8367623
-
项目类别:
-
资助金额:$39.16万
-
财政年份:2012
-
负责人:Nathan L Tintle
-
依托单位:
Analyzing the behavior and interpreting the results of gene based tests of rare variant association
-
批准号:9813293
-
项目类别:
-
资助金额:$35.27万
-
财政年份:2012
-
负责人:Nathan L Tintle
-
依托单位:
Evaluating the Cost Effectiveness of Alternative Sample Designs for Genetic Assoc
-
批准号:7841342
-
项目类别:
-
资助金额:$1.93万
-
财政年份:2009
-
负责人:Nathan L Tintle
-
依托单位:
Evaluating the Cost Effectiveness of Alternative Sample Designs for Genetic Assoc
-
批准号:8264409
-
项目类别:
-
资助金额:$7.48万
-
财政年份:2008
-
负责人:Nathan L Tintle
-
依托单位:
Evaluating the Cost Effectiveness of Alternative Sample Designs for Genetic Assoc
-
批准号:7363067
-
项目类别:
-
资助金额:$19.35万
-
财政年份:2008
-
负责人:Nathan L Tintle
-
依托单位:
海外基金