CAREER: Robust and scalable genome-wide phylogenetics
CAREER: Robust and scalable genome-wide phylogenetics
批准号:
1845967
负责人:
Siavash Mir arabbaygi
金额:
$54.92万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-02-15 至 2024-01-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The present diversity of life has evolved from a single ancestor through billions of years of evolution. Understanding these evolutionary histories is fascinating, but more importantly, is a crucial precursor to many biological analyses. Some evolutionary relationships are obvious (e.g., a cat is closer to a lion than a chicken) but other consequential relationships are hard to discern. Luckily, evolution operates on the genomes of organisms, and the sequence of genetic changes leaves a trace of the evolutionary histories. Following these traces and reconstructing the evolutionary past, however, is a computational problem, and as it turns out, is a difficult problem. Sophisticated methods are needed to infer a phylogeny: a tree, called tree-of-life, that shows the historical relationships between species. When sequencing whole genomes became possible in the mid-2000s, many believed the sheer amount of data would result in robust reconstructions of phylogenies. While genome sequencing has fulfilled some of its promises, other challenges remain. Large-scale data are hard to adequately model and are hard to screen for errors. As a result, different analyses do not always agree, and also, inference algorithms are pushed to their limits of scalability. Thus, an improved understanding of the tree-of-life requires not just more data but also better algorithms. Interestingly, as data sciences permeate many areas of science, issues of robustness to error and scalability faced in phylogenetics will confront many disciplines. Thus, the next generation of data scientists needs to be trained to consider these concerns when developing algorithms for data analysis.This project seeks to address current limitations in phylogenomics (phylogeny inference from whole genomes) and to integrate issues of robustness and scalability into teaching. The main challenge in phylogenomics is data heterogeneity, and there are two sources of data heterogeneity: real biological processes driving genome evolution that lead to discordant histories across the genome, and artefactual heterogeneity that results from complex pipelines used to prepare the data for inference. Models of real heterogeneity exist. However, current methods often require knowing the source of heterogeneity in advance, are often not scalable, are not always robust to artefactual heterogeneity. The approach taken here is to combine unsupervised learning and discrete optimization to build methods for identifying errors. These techniques will strive to minimize assumptions and will use both parametric and non-parametric statistics. The project will draw on machine learning, multi-criteria optimization, and high-performance computing. If successful, it will dramatically improve the accuracy and scalability of genome-wide phylogeny reconstruction and will help researchers understand intricate patterns in genome evolution. To integrate research and education, this project will enable yearly hackathons that bring together students with computational and biological expertise with the goal of developing robust and scalable methods. The project will also seek to improve the understanding of data science for undergrad and K-12 students, emphasizing for them both the excitement and challenges of analyzing large error-prone datasets. The tools developed here will be publicly available and well-documented. Yearly workshops will be held to help biologists learn and use the tools.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(22)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1371/journal.pone.0221068
发表时间:
2019-08-22
期刊:
PLOS ONE
影响因子:
3.7
作者:
[Balaban, Metin, Moshiri, Niema, Mirarab, Siavash]
通讯作者:
Mirarab, Siavash
SODA: multi-locus species delimitation using quartet frequencies
SODA:使用四重频率进行多位点物种界定
DOI:
10.1093/bioinformatics/btaa1010
发表时间:
2020
期刊:
Bioinformatics
影响因子:
5.8
作者:
[Rabiee, Maryam, Mirarab, Siavash]
通讯作者:
Mirarab, Siavash
DOI:
10.1093/bioinformatics/btab875
发表时间:
2022-01-03
期刊:
BIOINFORMATICS
影响因子:
5.8
作者:
[Mai, Uyen, Mirarab, Siavash]
通讯作者:
Mirarab, Siavash
Multispecies Coalescent: Theory and Applications in Phylogenetics
多物种合并:系统发育学的理论与应用
DOI:
10.1146/annurev-ecolsys-012121-095340
发表时间:
2021
期刊:
and Systematics
影响因子:
--
作者:
[Mirarab, Siavash, Nakhleh, Luay, Warnow, Tandy]
通讯作者:
Warnow, Tandy
TAPER: Pinpointing errors in multiple sequence alignments despite varying rates of evolution
TAPER:尽管进化速度不同,但仍可精确定位多个序列比对中的错误
DOI:
10.1111/2041-210x.13696
发表时间:
2021
期刊:
Methods in Ecology and Evolution
影响因子:
6.6
作者:
[Zhang, Chao, Zhao, Yiming, Braun, Edward L., Mirarab, Siavash]
通讯作者:
Mirarab, Siavash
共 14 条
III: Small: New algorithms for genome skimming and its applications
-
批准号:1815485
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2018
-
负责人:Siavash Mir arabbaygi
-
依托单位:
CRII: III: Using Genomic Context to Understand Evolutionary Histories of Individual Genes
-
批准号:1565862
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2016
-
负责人:Siavash Mir arabbaygi
-
依托单位:
国内基金
海外基金
登录
查看更多内容
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
-
批准号:70601028
-
项目类别:青年科学基金项目
-
资助金额:7.0万元
-
批准年份:2006
-
负责人:王明征
-
依托单位:
心理紧张和应力影响下Robust语音识别方法研究
-
批准号:60085001
-
项目类别:专项基金项目
-
资助金额:14.0万元
-
批准年份:2000
-
负责人:韩纪庆
-
依托单位:
ROBUST语音识别方法的研究
-
批准号:69075008
-
项目类别:面上项目
-
资助金额:3.5万元
-
批准年份:1990
-
负责人:高雨青
-
依托单位:
改进型ROBUST序贯检测技术
-
批准号:68671030
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1986
-
负责人:刘有恒
-
依托单位: