AF: Small: Fast and accurate computational tools for large-scale evolutionary inference: a phylogenetic network approach
AF: Small: Fast and accurate computational tools for large-scale evolutionary inference: a phylogenetic network approach
批准号:
1714417
负责人:
Kevin Liu
金额:
$40.47万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-15 至 2021-07-31
中文摘要
科学上的一个重大挑战是重建“生命之树”,这是地球上所有物种的系统发育或进化史。生命之树的概念反映了达尔文对进化的看法,即树状进化,其中分叉物种形成导致一个祖先物种产生两个基因隔离的后代物种。然而,最近的研究对这一观点提出了挑战。由于物种间的基因流动--即DNA在同时存在的物种之间共享--导致的“非树状”进化显著地塑造了比人们想象的更大的物种多样性的进化,包括人类和尼安德特人、老鼠和蝴蝶。在这些情况下,系统发育不能用树来描述,而是一个更一般的结构,称为系统发育网络。我们对进化论和生物学的理解正处于十字路口。在《生命之树》中,传统的树状进化假设被违反的频率有多高?基因流的进化作用是什么?应用包括了解抗生素耐药性在细菌中的传播,这每年给美国造成超过350亿美元的损失和2.3万人的生命损失,以及对杂草、老鼠和其他害虫的杀虫剂耐药性,这每年给美国造成数十亿美元的损失。系统发育学,或寻求利用生物分子序列和其他生物数据重建进化史的学科,可以为这些问题带来新的曙光。系统发育重建和分析需要两个要素:适用于所研究生物体的生物数据,以及能够有效和准确地分析数据的计算方法。今天,由于最近的生物技术进步,生物数据非常丰富,大规模的数据集也很常见。然而,计算方法并没有跟上步伐。在大数据时代,需要新的计算框架来快速、准确地进行系统发育网络推理和分析。为了应对这些挑战,该项目将利用大规模基因组序列数据集和连续生物数据的进化分析,为快速和准确的基于网络的系统发育推断创建新的计算框架。新的方法将通过一项全面的绩效研究加以验证。更广泛地说,这个项目将使学生培训、科学推广、开放源码软件开发和科学研究可能产生新的生物和生物医学发现。系统发育通常通过生物分子序列数据的计算分析来推断,系统发育比较方法用于连续生物数据(例如特征数据)的进化分析。如今,由于测序和相关生物技术的快速进步,“大数据”的挑战比比皆是。特别是,包含数百个基因组的大规模数据集现在很常见。因此,系统发育推断的现状面临着两个关键的可扩展性挑战:(1)研究中的生物数量,以及(2)更大的进化分歧,反映了树状和非树状进化的复杂相互作用。对于离散序列数据,最先进的方法解决了第二个挑战,但不能扩展到只有几十个基因组的输入;对于连续数据,需要可扩展的方法来解决系统发育不确定性和适应性进化背景下的第二个挑战。这项拟议的研究创造了新的计算方法,既解决了离散序列数据的挑战,也解决了连续数据的挑战。第一个目标是创建一个新的计算框架,用于使用大规模基因组序列数据进行可扩展的系统发育网络推理。该框架利用多物种网络合并模型来解释遗传漂移、不完全谱系排序和基因流动,以及传统的基于替换的序列进化模型。该框架建立在PI在大规模系统发育树推理方面的工作基础上,通过将分而治之的算法应用于更一般的网络情况,从而产生准确和高效的推理。第二个目标是发展新的随机模型和方法来分析系统发育网络上的连续特征进化。新的模型将推广广泛使用的假定树状进化的持续特征进化的非中性模型,并将被用于创建使用异质大规模输入进行系统发育推断的新方法。第三个目标是使用新的经验基准和合成基准来验证新的计算方法。经验基准包括通过持续合作产生的老鼠、植物和真菌数据集。
英文摘要
A grand challenge in science is reconstructing the "Tree of Life", which is the phylogeny, or evolutionary history, of all species on Earth. The notion of a Tree of Life reflects Darwin's view of evolution as "tree-like", where bifurcating speciation results in an ancestral species giving rise to two genetically isolated descendant species. However, recent studies have challenged this view. "Non-tree-like" evolution due to inter-species gene flow - where DNA is shared between species existing at the same time - has significantly shaped the evolution of a far greater diversity of species than was ever thought possible, including humans and Neandertal, mice, and butterflies. In these cases, the phylogeny cannot be described by a tree, but is instead a more general structure known as a phylogenetic network. Our understanding of evolution and biology is at a crossroads. How frequently is the traditional assumption of tree-like evolution violated in the Tree of Life? What is the evolutionary role of gene flow? Applications include understanding the spread of antibiotic resistance among bacteria, which costs the U.S. over $35 billion and a loss of 23,000 lives annually, and pesticide resistance in weeds, mice, and other pests, which costs the U.S. billions of dollars annually. Phylogenetics, or the discipline which seeks to reconstruct evolutionary histories using biomolecular sequences and other biological data, can shed new light into these questions. Two ingredients are necessary for phylogenetic reconstruction and analysis: suitable biological data for the organisms under study, and computational methods capable of efficiently and accurately analyzing the data. Today, biological data abounds due to recent biotechnological advances, and large-scale datasets are common. However, computational methods have not kept pace. New computational frameworks are needed for fast and accurate phylogenetic network inference and analysis in the era of "big data". To address these challenges, this project will create new computational frameworks for fast and accurate network-based phylogenetic inference using large-scale genomic sequence datasets and evolutionary analysis of continuous biological data. The new methodologies will be validated using a comprehensive performance study. More broadly, this project will enable student training, scientific outreach, open-source software development, and scientific research that may yield new biological and biomedical discoveries.Phylogenies are typically inferred using computational analysis of biomolecular sequence data, and phylogenetic comparative methods are used for evolutionary analysis of continuous biological data (e.g., trait data). Today, "big data" challenges abound due to rapid advances in sequencing and related biotechnologies. In particular, large-scale datasets with hundreds of genomes are now common. The state of the art of phylogenetic inference therefore faces two critical scalability challenges: (1) the number of organisms in a study, and (2) greater evolutionary divergence reflecting the complex interplay of tree-like and non-tree-like evolution. For discrete sequence data, state-of-the-art methods address the second challenge, but are not scalable beyond inputs with a few dozen genomes; for continuous data, scalable approaches are needed to address the second challenge in the context of phylogenetic uncertainty and adaptive evolution. The proposed research creates new computational approaches that address both challenges for discrete sequence data and continuous data. The first objective is to create a novel computational framework for scalable phylogenetic network inference using large-scale genomic sequence data. The framework makes use of the multi-species network coalescent model to account for genetic drift, incomplete lineage sorting, and gene flow as well as traditional substitution-based models of sequence evolution. The framework builds on the PI's work on large-scale phylogenetic tree inference by adapting divide-and-conquer algorithms to the more general case of networks, resulting in accurate and efficient inference. The second objective is to develop novel stochastic models and methods for analyzing continuous character evolution on phylogenetic networks. The new models will generalize widely-used non-neutral models of continuous character evolution that assume tree-like evolution, and will be used to create new methods for phylogenetic inference using heterogeneous large-scale inputs. The third objective is to validate the new computational methodologies using new empirical and synthetic benchmarks. The empirical benchmarks include mouse, plant, and fungal datasets that have been produced through ongoing collaborations.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Statistical analysis of GC-biased gene conversion and recombination hotspots in eukaryotic genomes: a phylogenetic hidden Markov model-based approach
真核基因组中GC偏向基因转换和重组热点的统计分析:基于系统发育隐马尔可夫模型的方法
DOI:
10.1145/3459930.3469509
发表时间:
2021
期刊:
Proceedings of ACM BCB 2021
影响因子:
--
作者:
[Gao, Meijun, Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
DOI:
10.1145/3107411.3107490
发表时间:
2017-08
期刊:
Proceedings of the 8th ACM International Conference on Bioinformatics, Computational Biology,and Health Informatics
影响因子:
--
作者:
[Hussein A. Hejase;N. V. Pol;G. Bonito;P. Edger;Kevin J. Liu]
通讯作者:
Hussein A. Hejase;N. V. Pol;G. Bonito;P. Edger;Kevin J. Liu
Non-parametric and semi-parametric support estimation using SEquential RESampling random walks on biomolecular sequences
使用生物分子序列上的 SEequential RESampling 随机游走进行非参数和半参数支持估计
DOI:
10.1186/s13015-020-00167-0
发表时间:
2020
期刊:
Algorithms for Molecular Biology
影响因子:
1
作者:
[Wang, Wei, Smith, Jack, Hejase, Hussein A., Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
An Application of Random Walk Resampling to Phylogenetic HMM Inference and Learning
随机游走重采样在系统发育 HMM 推理和学习中的应用
DOI:
10.1109/tnb.2020.2991302
发表时间:
2020
期刊:
IEEE Transactions on NanoBioscience
影响因子:
3.9
作者:
[Wang, Wei, Wuyun, Qiqige, Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
Scalable Statistical Introgression Mapping Using Approximate Coalescent-Based Inference
使用基于近似合并的推理的可扩展统计渗入映射
DOI:
10.1145/3307339.3343352
发表时间:
2019
期刊:
Computational Biology and Health Informatics
影响因子:
--
作者:
[Wuyun, Qiqige, VanKuren, Nicholas W., Kronforst, Marcus, Mullen, Sean P., Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
共 6 条
CAREER: Future phylogenies: novel computational frameworks for biomolecular sequence analysis involving complex evolutionary origins
-
批准号:2144121
-
项目类别:Continuing Grant
-
资助金额:$58.57万
-
财政年份:2022
-
负责人:Kevin Liu
-
依托单位:
CRII: AF: Novel evolutionary models and algorithms to connect genomic sequence and phenotypic data
-
批准号:1565719
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2016
-
负责人:Kevin Liu
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: