Computational Methods for Next-Generation GWAS
Computational Methods for Next-Generation GWAS
批准号:
9910009
负责人:
Christopher J Battey
金额:
$1.93万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-05-01 至 2020-07-31
关键词:
AgricultureBenchmarkingBiologyBreedingCommunitiesComputer softwareComputing MethodologiesCoupledCulicidaeDNA SequenceDataDimensionsEnvironmentEvolutionGene FrequencyGeneticGenotypeGeographic LocationsGoalsGuidelinesHaplotypesHealthHeart DiseasesHeightHumanImageLearningLinear ModelsLinear RegressionsLinkMachine LearningMeasuresMethodologyMethodsModelingNon-Insulin-Dependent Diabetes MellitusOligogenic TraitsOutputPerformancePhenotypePolygenic TraitsPopulationPopulation GeneticsPopulation HeterogeneityPositioning AttributeProcessPublic HealthRunningSamplingSignal TransductionSpatial DistributionStratificationStructureSumTechniquesTestingTrainingTrans-Omics for Precision MedicineVariantautoencoderbasebiobankcohortdeep learningdeep neural networkdiverse dataexperiencegenome wide association studygenome-widegenomic datahuman dataimage reconstructionimprovedlarge scale simulationlearning strategymachine learning algorithmneural networknext generationpolygenic risk scorepopulation stratificationsimulationstatisticssupervised learningtooltrait
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Project Summary/Abstract
Predicting phenotypes from DNA sequence variation is a major goal for genetics with potential
applications in evolutionary biology, crop breeding, and public health. A central challenge in this task is
separating genetic and environmental effects on phenotypes. In natural populations breeding structure is often
correlated with the environment across space such that different subpopulations experience different
environments. For genome-wide association studies (GWAS) this creates a problem: genetic and
environmental effects can be confounded by population structure, leading to inflated test statistics and low
predictive power across populations (Bulik-Sullivan et al. 2015, Mathieson and Mcvean, 2012). Understanding
when association studies are biased by population stratification and creating better methods to correct for it are
thus important challenges for population genetics over the next decade.
To identify conditions under which existing methods of population stratification correction are subject to
bias and develop robust new alternatives suitable for use with the continental-scale genomic datasets that are
now routinely available for humans, we propose to use simulations and machine learning to separate the
signals of fine-scale ancestry from polygenic phenotype association. In our first aim we will develop simulations
of polygenic phenotype evolution in continuous space and use the output to evaluate existing methods of
stratification control including linear mixed models, PC correction, and LD score regression. In this aim we will
seek to identify the regions of parameter space – i.e. the strength of isolation by distance and the spatial
distribution of environmental variation – in which existing methods can be expected to produce reliable effect
size estimates, and establish guidelines for applications of GWAS to structured populations.
We will then train machine learning algorithms on real genotype data from humans and mosquitoes to
describe continuous structure in large spatial samples using a variational autoencoder, a dimensionality
reduction technique based on deep neural networks that can take advantage of both allele frequency and
haplotype-based measures of differentiation in a single analysis and thus offer improved control of stratification
inflation in GWAS relative to the now standard PCA regression approach. Last we will apply deep learning
techniques to the problem of linking phenotypes and genotypes in structured samples by training neural
networks on simulated phenotypes and empirical genetic data. By training our networks on empirical genetic
data and incorporating contextual information about surrounding haplotype structure into the model, our
networks should learn to discriminate causal associations from false positives created by population structure
in the sample cohort, which will improve performance when attempting to identify associations with the real
phenotype. These methods will be applied to existing genomic datasets of height in humans, tested against the
current state-of-the-art approaches, and packaged as scalable software for the broader scientific community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
企业绩效评价的DEA-Benchmarking方法及动态博弈研究
-
批准号:70571028
-
项目类别:面上项目
-
资助金额:16.5万元
-
批准年份:2005
-
负责人:杨印生
-
依托单位: