AF: Small: Computational Methods for Large-scale Inference of Population History
AF: Small: Computational Methods for Large-scale Inference of Population History
批准号:
1718093
负责人:
Yufeng Wu
金额:
$40.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2022-08-31
中文摘要
考虑来自一个种群的几个无关的人类个体。这是一个常识,这些人是一些共同祖先的后代,如果追溯到足够长的时间。事实上,这些样本个体共享一个家谱历史,指定这些个体的祖先。这个家谱是非常翔实的,因为它可以告诉,例如,哪些人是密切相关的。系谱学的一个潜在应用是理解为什么有些人比其他人更容易受到某些表型特征(如糖尿病或癌症)的影响。可能的情况是,共享一个特征的个体在家谱上的关系比其他群体更密切。系谱学虽然有用,但不能直接观察。所谓的遗传变异提供了样本个体潜在谱系的线索。一种常见的遗传变异类型是单核苷酸多态性(SNP)。SNP是指群体中的个体在该位置处可具有不同核苷酸的基因组位置。思考一下,在SNP上具有相同核苷酸的个体往往比具有不同核苷酸的个体更接近。这可以允许人们从在多个SNP处收集的遗传数据推断合理的潜在系谱。然而,从真实的SNP数据推断系谱比这复杂得多。一个主要的困难是由减数分裂重组引起的。如果没有重组,家谱可以被建模为一棵树(类似于生物学中广泛研究的生命树模型)。简化允许一个基因组有一个以上的祖先,因此违反了这个简单的树模型的基本属性。进化基本上将基因组分解成许多小片段,其中每个片段可能起源于不同的祖先。也就是说,不同基因组位置的系谱历史可能不同。因此,有重组的系谱学比没有重组的系谱学复杂得多。本项目旨在开发有效的计算方法,用于分析过去几年中获得的大规模群体遗传数据。其主要目标是首先从遗传数据中准确推断样本个体的系谱历史,然后利用推断的系谱对若干群体遗传问题进行推断。拟议研究的成功完成将产生新的计算工具和软件,使人口遗传学家能够更好地了解大规模人口遗传数据的影响。这些工具的潜在应用包括,例如,绘制与复杂性状相关的基因组位置图,推断种群混合历史,以及发现处于自然选择下的基因组区域。这个项目将开发有效和准确的计算方法,用于根据推断的基因谱系从单倍型推断种群历史。基因系谱是指现存群体单倍型的进化历史,并捕获潜在的LD信息。虽然基因谱系是群体遗传学的基础,但大多数现有的推理方法都没有明确使用基因谱系,因为谱系不能直接观察到。由于基因组技术和系谱推断方法的最新发展,从单倍型推断基因系谱才刚刚开始变得可行。本项目旨在为以下两个问题开发有效的计算方法。首先,将开发从单倍型推断基因谱系的新方法。第二,将开发推断人口统计学历史(例如人口混合)的新方法。成功完成拟议的研究将产生新的高效和准确的算法,在实用的软件工具中实现,并允许人口生物学家从基因组规模的数据推断人口的历史。开发的软件工具将免费提供给多学科研究界,并有望在复杂的人口历史推断中实现新的生物学应用。研究成果将融入课堂教学。该项目将确保广泛传播研究成果和教材。拟议的教育和外联活动包括培训具有独特跨学科技能的未来研究人员。
英文摘要
Consider several unrelated human individuals from a population. It is a common sense that these individuals are descendants of some common ancestors if tracing backward in time long enough. Indeed, these sampled individuals share a genealogical history that specifies the ancestry of these individuals. This genealogy is very informative, since it can tell, e.g. which individuals are closely related. One potential application of the genealogy is understanding why some individuals are more susceptible to some phenotypic traits (such as diabetes or cancer) than others. It might be the case that individuals sharing a trait are more closely related to each other on the genealogy than the rest of the population. Genealogy, although useful, cannot be directly observed. The so-called genetic variation provides hints on the underlying genealogy of the sampled individuals. One common type of genetic variations is the single nucleotide polymorphism (SNP). A SNP refers to the genomic position where individuals in a population may have different nucleotides at the position. A moment of thoughts suggests that individuals with the same nucleotide at a SNP tend to be more closely related than individuals with different nucleotides. This may allow one to infer the plausible underlying genealogy from genetic data collected at multiple SNPs. Inference of genealogy from real SNP data is, however, much more complex than this. One main difficulty is caused by meiotic recombination. Without recombination, genealogy can be modeled as a tree (similar to the usual tree of life model that has been extensively studied in biology). Recombination allows one genome to have more than one ancestor and thus violates the basic property of this simple tree model. Recombination essentially breaks down the genome into many small segments, where each segment may originate from different ancestors. That is, genealogical history at different genomic positions may be different. Genealogy with recombination is thus much more complex than that with no recombination.This project aims to developing effective computational methods for analyzing large-scale population genetic data that has become available during the past several years. The main goals are first accurately inferring the genealogical history of sampled individuals from the genetic data, and then performing inference for several population genetic problems with the inferred genealogy. The successful completion of the proposed research will produce new computational tools and software that may allow population geneticists to better understand the implications of large-scale population genetic data. Potential applications of these tools include, for example, mapping the genomic positions that are associated with complex traits, inferring the population admixture history and finding regions of the genome that are under natural selection.The intellectual merits of this project are as follows. This project will develop efficient and accurate computational methods for inferring population history from haplotypes based on inferred gene genealogies. Gene genealogy refers to the evolutionary history of extant population haplotypes, and captures the underlying LD information. While gene genealogies are fundamental to population genetics, most existing inference methods don't use gene genealogies explicitly because genealogies are not directly observable. Inferring gene genealogies from haplotypes is just starting to become feasible, due to the latest development in genomic technologies and genealogy inference methods. This project aims to developing effective computational methods for the following two problems. First, new methods for inferring gene genealogies from haplotypes will be developed. Second, new methods for inferring population demographic history (e.g. population admixture) will be developed. Successful completion of the proposed research will produce new efficient and accurate algorithms that are implemented in practical software tools and allow population biologists to infer population history from genome-scale data.The broader impacts of this project include the following. Developed software tools will be made available freely to the multidisciplinary research community, and are expected to enable novel biological applications in complex population history inference. Research results will be integrated into classroom teaching. The project will ensure broad dissemination of the research results and teaching materials. The proposed educational and outreach activities include training of future researchers with unique interdisciplinary skills.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1093/bioinformatics/btaa465
发表时间:
2020-07-01
期刊:
BIOINFORMATICS
影响因子:
5.8
作者:
[Wu, Yufeng]
通讯作者:
Wu, Yufeng
DOI:
10.1371/journal.pcbi.1008065
发表时间:
2020-08-01
期刊:
PLOS COMPUTATIONAL BIOLOGY
影响因子:
4.3
作者:
[Pei, Jingwen, Zhang, Yiming, Wu, Yufeng]
通讯作者:
Wu, Yufeng
CLADES: A classification-based machine learning method for species delimitation from population genetic data
CLADES:一种基于分类的机器学习方法,用于从种群遗传数据中进行物种界定
DOI:
10.1111/1755-0998.12887
发表时间:
2018
期刊:
Molecular Ecology Resources
影响因子:
7.7
作者:
[Pei, Jingwen, Chu, Chong, Li, Xin, Lu, Bin, Wu, Yufeng]
通讯作者:
Wu, Yufeng
DOI:
10.1186/s12864-019-6154-7
发表时间:
2018-04
期刊:
BMC Genomics
影响因子:
4.4
作者:
[Xin Li;Yufeng Wu]
通讯作者:
Xin Li;Yufeng Wu
DOI:
10.1038/s41598-020-71300-7
发表时间:
2020-08-31
期刊:
SCIENTIFIC REPORTS
影响因子:
4.6
作者:
[Inkman, Matthew J., Jayachandran, Kay, Zhang, Jin]
通讯作者:
Zhang, Jin
共 6 条
III: Small: Computational Methods for Ancestry Inference In Genetics
-
批准号:1909425
-
项目类别:Standard Grant
-
资助金额:$41.11万
-
财政年份:2019
-
负责人:Yufeng Wu
-
依托单位:
III: Small: Computational Methods for Analyzing Complex Genomes with Sequence Data
-
批准号:1526415
-
项目类别:Standard Grant
-
资助金额:$42.63万
-
财政年份:2015
-
负责人:Yufeng Wu
-
依托单位:
AF: Small: Algorithms for Reconstructing Complex Evolutionary History with Discordant Phylogenetic Trees
-
批准号:1116175
-
项目类别:Standard Grant
-
资助金额:$25.68万
-
财政年份:2011
-
负责人:Yufeng Wu
-
依托单位:
CAREER: Efficient and Accurate Computation for High Throughput Sequencing Related Problems in Population Genomics
-
批准号:0953563
-
项目类别:Continuing Grant
-
资助金额:$49.64万
-
财政年份:2010
-
负责人:Yufeng Wu
-
依托单位:
III-CXT-Medium: Collaborative Research: Inference of Complex Genealogical Histories in Populations: Algorithms and Applications
-
批准号:0803440
-
项目类别:Standard Grant
-
资助金额:$30.52万
-
财政年份:2008
-
负责人:Yufeng Wu
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: