Deep learning for population genetics
Deep learning for population genetics
批准号:
9976348
负责人:
ANDREW D KERN
金额:
$52.92万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-04-21 至 2024-02-28
关键词:
AlgorithmsAreaBiologicalBiologyClassificationCodeCommunitiesComputer Vision SystemsComputer softwareDNA SequenceDataDevelopmentEnsureFloodsGeneticGenetic RecombinationGenomeGenomicsGenotypeGoalsImageLeadLearningLeftMachine LearningMeasuresMedicineMethodologyMethodsModelingModernizationNatural Language ProcessingNatural SelectionsNaturePerformancePopulationPopulation ExplosionsPopulation GeneticsProcessProgram DevelopmentPublishingResearch PersonnelSequence AlignmentSoftware ToolsTechniquesTechnologyTrainingTreesUncertaintyUrsidae FamilyWorkbasecomputational chemistryconvolutional neural networkdeep learningdeep neural networkdesignempoweredflexibilitygenetic informationgenome sequencinggenomic datainfancyinnovationlearning classifierlearning strategymachine learning algorithmmachine learning methodnetwork architectureneural networkneural network architecturenext generationnovelopen sourcerandom forestrecurrent neural networkresearch and developmentspeech recognitionstatisticsstemsuccesssupervised learningtooltool developmentuser friendly software
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Project Summary
The revolution in genome sequencing technologies over the past 15 years has created an explosion of
population genomic data but has left in its wake a gap in our ability to make sense of data at this scale. In
particular, whereas population genetics as a field has been traditionally data-limited, the massive volume of
current sequencing means that previously unanswerable questions may now be within reach. To capitalize on
this flood of information we need new methods and modes of analysis.
In the past 5 years the world of machine learning has been revolutionized by the rise of deep neural
networks. These so-called deep learning methods offer incredible flexibility as well as astounding
improvements in performance for a wide array of machine learning tasks, including computer vision, speech
recognition, and natural language processing. This proposal aims to harness the great potential of deep
learning for population genetic inference.
In recent years our group has made great strides in using supervised machine learning for population
genomic analysis (reviewed in Schrider and Kern 2018). However, this work has focused primarily on using
more traditional machine learning methods such as random forests. As we argue in this proposal, DNA
sequence data are particularly well suited for modern deep learning techniques, and we demonstrate that the
application of these methods can rapidly lead to state-of-the-art performance in very difficult population genetic
tasks such as estimating rates of recombination. The power of these methods for handling genetic data stems
in part from their ability to automatically learn to extract as much useful information as possible from an
alignment of DNA sequences in order to solve the task at hand, rather than relying on one or more predefined
summary statistics which are generally problem-specific and may omit information present in the raw data.
In this proposal we lay out a systematic approach for both empowering the field with these tools and
understanding their shortcomings. In particular, we propose to design deep neural networks for solving
population genetic problems, and incorporate successful networks into user-friendly software tools that will be
shared with the community. We will also investigate a variety of methods for estimating the uncertainty of
predictions produced by deep learning methods; this area is understudied in machine learning but of great
importance to biological researchers who require an accurate measure of the degree of uncertainty
surrounding an estimate. Finally, we will explore the impact of training data misspecification—wherein the data
used to train a machine learning method differ systematically from the data to which it will be applied in
practice. We will devise techniques to mitigate the impact of such misspecification in order to ensure that our
tools will be robust to the complications inherent in analyzing real genomic data sets. Together, these
advances have the potential to transform the methodological landscape of population genetic inference.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational Population Genetics
-
批准号:10552275
-
项目类别:
-
资助金额:$43.99万
-
财政年份:2023
-
负责人:ANDREW D KERN
-
依托单位:
Deep learning for population genetics
-
批准号:10349557
-
项目类别:
-
资助金额:$42.04万
-
财政年份:2020
-
负责人:ANDREW D KERN
-
依托单位:
Deep learning for population genetics
-
批准号:10574510
-
项目类别:
-
资助金额:$42.04万
-
财政年份:2020
-
负责人:ANDREW D KERN
-
依托单位:
Population genomics of adaptation
-
批准号:9383198
-
项目类别:
-
资助金额:$29.55万
-
财政年份:2017
-
负责人:ANDREW D KERN
-
依托单位:
POPULATION GENOMICS OF ADAPTATION
-
批准号:9753261
-
项目类别:
-
资助金额:$29.5万
-
财政年份:2017
-
负责人:ANDREW D KERN
-
依托单位:
Human Population Genomics
-
批准号:7053104
-
项目类别:
-
资助金额:$4.21万
-
财政年份:2005
-
负责人:ANDREW D KERN
-
依托单位:
Human Population Genomics
-
批准号:7283831
-
项目类别:
-
资助金额:$4.83万
-
财政年份:2005
-
负责人:ANDREW D KERN
-
依托单位:
Human Population Genomics
-
批准号:7146707
-
项目类别:
-
资助金额:$4.55万
-
财政年份:2005
-
负责人:ANDREW D KERN
-
依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
-
批准号:2021JJ40433
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:孙磊
-
依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
-
批准号:32001603
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:段真珍
-
依托单位:
AREA国际经济模型的移植.改进和应用
-
批准号:18870435
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1988
-
负责人:史树中
-
依托单位: