CAREER: Future phylogenies: novel computational frameworks for biomolecular sequence analysis involving complex evolutionary origins
CAREER: Future phylogenies: novel computational frameworks for biomolecular sequence analysis involving complex evolutionary origins
批准号:
2144121
负责人:
Kevin Liu
金额:
$58.57万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-03-01 至 2027-02-28
中文摘要
该奖项的全部或部分资金来自《2021年美国救援计划法案》(公法117-2)。系统发育学是一门试图重建和分析一组生物体的系统发育或进化史的学科。系统发育重建主要通过对DNA和其他生物分子序列数据的计算分析来完成。系统发育及其提供的进化见解对生物学和其他学科以及许多应用至关重要:重要的例子包括重建和研究生命树--地球上所有生命的进化史、了解人类起源、传染病流行病学和发现未来流行病的新解决方案、作物改良和农业以及法医学。由于生物分子测序技术的进步,系统发育研究所需的两个关键因素之一出现了重大飞跃:现有生物分子数据的规模现在是任何领域中最大的之一,到2025年,生物分子数据的速度和存储预计将与Twitter和YouTube相媲美或更大。另一方面,最近的大数据系统发育研究指出,关于系统发育的两个关键组成部分中的第二个,存在一个关键差距:现有的计算算法需要超越其关于生物分子序列进化的传统简化假设。其中两个最重要的假设是:(1)忽略生物分子序列内在顺序性质的“不知道序列”的方法,以及(2)预先提出的假设,即进化关系具有简单的分支结构并呈“树状”--即可以用树或其他简单的表示来准确地描述。需要新的计算方法和基础设施来超越这些传统的假设,并开启“未来系统发生学”和下一代系统发生学的研究。因此,该项目将为生物分子序列数据的复杂系统发育分析创造新的开创性模型和算法。该项目还通过开发新的课程和与位于密歇根州中部的儿童科学博物馆Impression 5 Science Center合作,解决STEM教育方面的差距。将通过开放源码软件分发和开放数据资源、已开发的软件和数据基础设施带来的新的科学发现、科学推广活动以及强调多样性、公平和包容性(DEI)的学生培训和辅导来扩大项目的影响。该项目将推动计算系统发育领域的多个前沿。第一个研究目标是开发新的统计重采样算法,超越生物分子数据被假定为独立和相同分布(I.I.D.)的“无信息”分析,转向“有信息”的序列感知分析;一个中心方法将是利用机器学习的最新进展。新的算法将被用来更好地评估系统发育分析和其他关键路径分析任务的严密性和重复性。第二个研究目标是创建数学理论、统计模型和计算算法,以超越传统的系统发育表示法(例如,系统发育树等),转向更一般的复杂基因组进化的图论模型。第三个研究目标是对前两个研究目标的计算框架进行全面验证和业绩评估研究。这些研究将利用合成和经验基准数据集,这些数据集捕获了广泛的进化条件和数据集特征。该项目还包括两个教育目标:一门关于跨学科计算机科学的新的环境和信息技术主题的课程,以及一项将在印象5科学中心展出的关于技术和计算机编程的新博物馆展览。开放源码软件和开放数据成果将推动未来的方法学研究,并使其他方面无法获得的科学发现成为可能,而科学宣传将有助于播种和推动对该项目贡献的吸收。该项目还包括本科生和研究生的学生培训和辅导活动。项目交付成果和其他结果可以在https://gitlab.msu.edu/liulab.This上找到,该奖项反映了国家科学基金会的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).Phylogenetics is the discipline that seeks to reconstruct and analyze the phylogeny, or evolutionary history, of a set of organisms. Phylogenetic reconstruction is primarily accomplished through computational analysis of DNA and other biomolecular sequence data. Phylogenies and the evolutionary insights that they provide are essential to biology and other disciplines, as well as many applications: important examples include reconstructing and studying the Tree of Life - the evolutionary history of all life on Earth, understanding human origins, infectious disease epidemiology and discovery of new solutions to future pandemics, crop improvement and agriculture, and forensic science. One of the two key ingredients needed for phylogenetic studies has seen a major leap forward thanks to advances in biomolecular sequencing technology: the scale of available biomolecular data is now among the largest in any domain and, in 2025, biomolecular data velocity and storage is projected to be comparable to or larger than Twitter and YouTube. On the other hand, recent "big data" phylogenetic studies point to a critical gap regarding the second of the two key ingredients in phylogenetics: existing computational algorithms need to move beyond their traditional simplifying assumptions about biomolecular sequence evolution. Two of the most important assumptions are: (1) "sequence-unaware" methods that ignore the inherently sequential nature of biomolecular sequences, and (2) the pre hoc assumption that evolutionary relationships have a simple branching structure and are "tree-like" - i.e., can be accurately described by a tree or other simple representation. New computational approaches and infrastructure are needed to move beyond these traditional assumptions and unlock the study of "future phylogenies" and next-generation phylogenetics. This project will therefore create new pathbreaking models and algorithms for complex phylogenetic analyses of biomolecular sequence data. The project also addresses gaps in STEM education through new curriculum development and a collaboration with the Impression 5 Science Center, a children’s science museum in mid-Michigan. Project impacts will be broadened through open-source software distributions and open data resources, new scientific discoveries enabled by the developed software and data infrastructure, scientific outreach activities, and student training and mentoring with a strong emphasis on diversity, equity, and inclusion (DEI).This project will advance the field of computational phylogenetics along multiple frontiers. The first research objective is to develop new statistical resampling algorithms that move beyond "uninformed" analysis where biomolecular data are assumed to be independent and identically distributed (i.i.d.), and towards "informed" sequence-aware analysis; a central approach will be to make use of the latest advances in machine learning. The new algorithms will be used to better assess rigor and reproducibility during phylogenetic analyses and other critical-path analytical tasks. The second research objective is to create mathematical theory, statistical models, and computational algorithms to move beyond traditional phylogenetic representations (e.g., phylogenetic trees, etc.), and towards more general graph-theoretic models of complex genome evolution. The third research objective is to conduct comprehensive validation and performance assessment studies of the first two research objectives’ computational frameworks. The studies will utilize both synthetic and empirical benchmarking datasets that capture a wide range of evolutionary conditions and dataset features. The project also includes two educational objectives: a new course on DEI topics in interdisciplinary computer science, and a new museum exhibit on technology and computer programming that will be exhibited at the Impression 5 Science Center. Open-source software and open data deliverables will drive future methodological research and enable otherwise inaccessible scientific discoveries, and scientific outreach will help seed and drive uptake of the project’s contributions. The project also includes student training and mentoring activities at the undergraduate and graduate levels. Project deliverables and other results can be found at https://gitlab.msu.edu/liulab.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
The impact of gene sequence alignment and gene tree estimation error on summary-based species network estimation
基因序列比对和基因树估计误差对基于摘要的物种网络估计的影响
DOI:
10.1145/3535508.3545559
发表时间:
2022
期刊:
and Health Informatics (BCB ’22
影响因子:
--
作者:
[Gao, Meijun, Wang, Wei, Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
The Impact of Species Tree Estimation Error on Cophylogenetic Reconstruction
物种树估计误差对共系统发育重建的影响
DOI:
10.1145/3584371.3612964
发表时间:
2023
期刊:
and Health Informatics
影响因子:
--
作者:
[Zheng, Julia, Nishida, Yuya, Okrasinska, Alicja, Bonito, Gregory M., Heath-Heckman, Elizabeth A., Liu, Kevin J.]
通讯作者:
Liu, Kevin J.
Reconstructing Phylogenies Using Branch-Variable Substitution Models and Unaligned Biomolecular Sequences: A Performance Study and New Resampling Method
使用分支变量替换模型和未对齐的生物分子序列重建系统发育:性能研究和新的重采样方法
DOI:
10.1145/3584371.3613011
发表时间:
2023
期刊:
and Health Informatics
影响因子:
--
作者:
[Doko, Rei, Liu, Kevin]
通讯作者:
Liu, Kevin
AF: Small: Fast and accurate computational tools for large-scale evolutionary inference: a phylogenetic network approach
-
批准号:1714417
-
项目类别:Standard Grant
-
资助金额:$40.47万
-
财政年份:2017
-
负责人:Kevin Liu
-
依托单位:
CRII: AF: Novel evolutionary models and algorithms to connect genomic sequence and phenotypic data
-
批准号:1565719
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2016
-
负责人:Kevin Liu
-
依托单位:
海外基金