AF: Small: Fast and accurate computational tools for large-scale evolutionary inference: a phylogenetic network approach
AF: Small: Fast and accurate computational tools for large-scale evolutionary inference: a phylogenetic network approach
批准号:
1714417
负责人:
Kevin Liu
金额:
$40.47万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-15 至 2021-07-31
中文摘要
科学界面临的一个重大挑战是重建“生命之树”,即地球上所有物种的系统发育或进化史。生命之树的概念反映了达尔文对进化的“树状”观点,即物种形成的分叉导致一个祖先物种产生两个基因隔离的后代物种。然而,最近的研究挑战了这一观点。由于物种间基因流动而产生的“非树状”进化——即同时存在的物种之间共享DNA——显著地塑造了物种多样性的进化,其多样性远远超出了人们的想象,包括人类和尼安德特人、老鼠和蝴蝶。在这些情况下,系统发育不能用树来描述,而是一种更一般的结构,称为系统发育网络。我们对进化和生物学的理解正处在一个十字路口。传统的树状进化假设在生命之树中被违反的频率有多高?基因流的进化作用是什么?应用包括了解细菌之间抗生素耐药性的传播,这使美国每年损失超过350亿美元,并损失23,000人的生命,以及杂草,老鼠和其他害虫的农药耐药性,这使美国每年损失数十亿美元。系统遗传学,或者是一门试图利用生物分子序列和其他生物学数据重建进化历史的学科,可以为这些问题提供新的线索。系统发育重建和分析需要两个要素:适合所研究生物的生物学数据,以及能够有效和准确分析数据的计算方法。今天,由于最近的生物技术进步,生物数据丰富,大规模数据集很常见。然而,计算方法并没有跟上时代的步伐。在“大数据”时代,需要新的计算框架来进行快速准确的系统发育网络推理和分析。为了应对这些挑战,该项目将创建新的计算框架,利用大规模基因组序列数据集和连续生物数据的进化分析,快速准确地进行基于网络的系统发育推断。新方法将通过综合绩效研究得到验证。更广泛地说,该项目将使学生培训、科学推广、开源软件开发和可能产生新的生物和生物医学发现的科学研究成为可能。系统发生通常使用生物分子序列数据的计算分析来推断,系统发生比较方法用于连续生物数据(例如,性状数据)的进化分析。如今,由于测序和相关生物技术的快速发展,“大数据”挑战比比皆是。特别是,拥有数百个基因组的大规模数据集现在很常见。因此,系统发育推断的现状面临着两个关键的可扩展性挑战:(1)研究中的生物数量;(2)更大的进化分歧反映了树状和非树状进化的复杂相互作用。对于离散序列数据,最先进的方法解决了第二个挑战,但不能扩展到超过几十个基因组的输入;对于连续数据,需要可扩展的方法来解决系统发育不确定性和适应性进化背景下的第二个挑战。提出的研究创造了新的计算方法来解决离散序列数据和连续数据的挑战。第一个目标是创建一个新的计算框架,用于使用大规模基因组序列数据进行可扩展的系统发育网络推断。该框架利用多物种网络聚结模型来解释遗传漂变、不完全谱系分类和基因流以及传统的基于替代的序列进化模型。该框架建立在PI的大规模系统发育树推断工作的基础上,通过将分而治之算法应用于更一般的网络情况,从而产生准确和有效的推断。第二个目标是开发新的随机模型和方法来分析系统发育网络上的连续特征进化。新模型将推广广泛使用的假设树状进化的连续特征进化的非中性模型,并将用于创建使用异构大规模输入的系统发育推断的新方法。第三个目标是使用新的经验和合成基准来验证新的计算方法。经验基准包括通过正在进行的合作产生的小鼠、植物和真菌数据集。
英文摘要
A grand challenge in science is reconstructing the "Tree of Life", which is the phylogeny, or evolutionary history, of all species on Earth. The notion of a Tree of Life reflects Darwin's view of evolution as "tree-like", where bifurcating speciation results in an ancestral species giving rise to two genetically isolated descendant species. However, recent studies have challenged this view. "Non-tree-like" evolution due to inter-species gene flow - where DNA is shared between species existing at the same time - has significantly shaped the evolution of a far greater diversity of species than was ever thought possible, including humans and Neandertal, mice, and butterflies. In these cases, the phylogeny cannot be described by a tree, but is instead a more general structure known as a phylogenetic network. Our understanding of evolution and biology is at a crossroads. How frequently is the traditional assumption of tree-like evolution violated in the Tree of Life? What is the evolutionary role of gene flow? Applications include understanding the spread of antibiotic resistance among bacteria, which costs the U.S. over $35 billion and a loss of 23,000 lives annually, and pesticide resistance in weeds, mice, and other pests, which costs the U.S. billions of dollars annually. Phylogenetics, or the discipline which seeks to reconstruct evolutionary histories using biomolecular sequences and other biological data, can shed new light into these questions. Two ingredients are necessary for phylogenetic reconstruction and analysis: suitable biological data for the organisms under study, and computational methods capable of efficiently and accurately analyzing the data. Today, biological data abounds due to recent biotechnological advances, and large-scale datasets are common. However, computational methods have not kept pace. New computational frameworks are needed for fast and accurate phylogenetic network inference and analysis in the era of "big data". To address these challenges, this project will create new computational frameworks for fast and accurate network-based phylogenetic inference using large-scale genomic sequence datasets and evolutionary analysis of continuous biological data. The new methodologies will be validated using a comprehensive performance study. More broadly, this project will enable student training, scientific outreach, open-source software development, and scientific research that may yield new biological and biomedical discoveries.Phylogenies are typically inferred using computational analysis of biomolecular sequence data, and phylogenetic comparative methods are used for evolutionary analysis of continuous biological data (e.g., trait data). Today, "big data" challenges abound due to rapid advances in sequencing and related biotechnologies. In particular, large-scale datasets with hundreds of genomes are now common. The state of the art of phylogenetic inference therefore faces two critical scalability challenges: (1) the number of organisms in a study, and (2) greater evolutionary divergence reflecting the complex interplay of tree-like and non-tree-like evolution. For discrete sequence data, state-of-the-art methods address the second challenge, but are not scalable beyond inputs with a few dozen genomes; for continuous data, scalable approaches are needed to address the second challenge in the context of phylogenetic uncertainty and adaptive evolution. The proposed research creates new computational approaches that address both challenges for discrete sequence data and continuous data. The first objective is to create a novel computational framework for scalable phylogenetic network inference using large-scale genomic sequence data. The framework makes use of the multi-species network coalescent model to account for genetic drift, incomplete lineage sorting, and gene flow as well as traditional substitution-based models of sequence evolution. The framework builds on the PI's work on large-scale phylogenetic tree inference by adapting divide-and-conquer algorithms to the more general case of networks, resulting in accurate and efficient inference. The second objective is to develop novel stochastic models and methods for analyzing continuous character evolution on phylogenetic networks. The new models will generalize widely-used non-neutral models of continuous character evolution that assume tree-like evolution, and will be used to create new methods for phylogenetic inference using heterogeneous large-scale inputs. The third objective is to validate the new computational methodologies using new empirical and synthetic benchmarks. The empirical benchmarks include mouse, plant, and fungal datasets that have been produced through ongoing collaborations.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Statistical analysis of GC-biased gene conversion and recombination hotspots in eukaryotic genomes: a phylogenetic hidden Markov model-based approach
真核基因组中GC偏向基因转换和重组热点的统计分析:基于系统发育隐马尔可夫模型的方法
DOI:
10.1145/3459930.3469509
发表时间:
2021
期刊:
Proceedings of ACM BCB 2021
影响因子:
--
作者:
[Gao, Meijun, Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
DOI:
10.1145/3107411.3107490
发表时间:
2017-08
期刊:
Proceedings of the 8th ACM International Conference on Bioinformatics, Computational Biology,and Health Informatics
影响因子:
--
作者:
[Hussein A. Hejase;N. V. Pol;G. Bonito;P. Edger;Kevin J. Liu]
通讯作者:
Hussein A. Hejase;N. V. Pol;G. Bonito;P. Edger;Kevin J. Liu
Non-parametric and semi-parametric support estimation using SEquential RESampling random walks on biomolecular sequences
使用生物分子序列上的 SEequential RESampling 随机游走进行非参数和半参数支持估计
DOI:
10.1186/s13015-020-00167-0
发表时间:
2020
期刊:
Algorithms for Molecular Biology
影响因子:
1
作者:
[Wang, Wei, Smith, Jack, Hejase, Hussein A., Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
An Application of Random Walk Resampling to Phylogenetic HMM Inference and Learning
随机游走重采样在系统发育 HMM 推理和学习中的应用
DOI:
10.1109/tnb.2020.2991302
发表时间:
2020
期刊:
IEEE Transactions on NanoBioscience
影响因子:
3.9
作者:
[Wang, Wei, Wuyun, Qiqige, Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
Scalable Statistical Introgression Mapping Using Approximate Coalescent-Based Inference
使用基于近似合并的推理的可扩展统计渗入映射
DOI:
10.1145/3307339.3343352
发表时间:
2019
期刊:
Computational Biology and Health Informatics
影响因子:
--
作者:
[Wuyun, Qiqige, VanKuren, Nicholas W., Kronforst, Marcus, Mullen, Sean P., Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
共 6 条
CAREER: Future phylogenies: novel computational frameworks for biomolecular sequence analysis involving complex evolutionary origins
-
批准号:2144121
-
项目类别:Continuing Grant
-
资助金额:$58.57万
-
财政年份:2022
-
负责人:Kevin Liu
-
依托单位:
CRII: AF: Novel evolutionary models and algorithms to connect genomic sequence and phenotypic data
-
批准号:1565719
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2016
-
负责人:Kevin Liu
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: