CAREER: New Statistical Approaches for Studying Evolutionary Processes: Inference, Attribution and Computation
CAREER: New Statistical Approaches for Studying Evolutionary Processes: Inference, Attribution and Computation
批准号:
2143242
负责人:
Julia Palacios
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-02-01 至 2027-01-31
中文摘要
该奖项全部或部分根据2021年美国救援计划法案(公法117-2)资助。从DNA等分子序列样本进行统计推断提出了一系列基本挑战。这些挑战包括对样本的祖先和过去的进化历史进行复杂的建模,以及大量的噪声数据。遗传数据的持续大规模增加导致目前的方法不适用于现有的数据量,研究人员被迫减少现有数据的样本或从不充分的汇总统计数据中推断参数。 这个研究项目将解决从现代分子数据推断的最佳设计的聚结建模的需要。聚结是一个关于系谱的概率模型,也就是说,代表样本祖先的树。聚结模型用于推断与科学相关的参数,如有效种群规模、迁移模式和选择。该项目的研究目标是扩大聚结模型的类别,并设计新的高效统计算法,使我们能够解决许多推动科学发展的实际问题。此外,这些项目的成果将促进新的统计理论和易于处理的方法的发展,有助于生物解决方案。该项目还概述了一项积极的计划,以开展广泛的教育和外联活动,扩大对统计科学的参与,并增强科学领域更具包容性的氛围。参与该项目的本科生和研究生将在统计科学和生物学的界面上获得跨学科实践研究培训的独特机会,使他们能够为进化生物学,分子生物学,群体遗传学,遗传学,癌症基因组学,概率建模,统计推断和相关领域的进展做出贡献。PI将积极参与多项外展活动,例如斯坦福大学的本科生暑期研究计划,该计划将允许招募更多不同的未来数据科学家,并在科学领域培养更具包容性的气候。 该项目的研究成果将作为统计遗传学新课程的基础,并将纳入本科和研究生课程。 具体来说,这个项目将扩大类的合并模型,并提供一套新的算法和统计方法,利用度量概念的系谱,集总的马尔可夫链和分而治之的战略。具体目标包括:(1)开发结合模型,以纳入各种采样方案和生物过程,如动态种群结构,重组和强选择;(2)开发结合理论和应用的度量框架;(3)为进化参数的贝叶斯推断开发可扩展的策略,以及(4)实现,验证和分析传染病的分子序列,如SARS-CoV-2,古代和现代人类DNA样本以及癌症单细胞变异。此外,该项目将积极促进在多个方面扩大对统计科学的参与,从基于团队的跨学科研究培训到社区外展。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的知识价值和更广泛的影响审查标准进行评估来支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2). Statistical inference from a sample of molecular sequences such as DNA poses a series of fundamental challenges. These challenges include complex modeling of the sample's ancestry and past evolutionary history, large and noisy data. The ongoing large-scale increase of genetic data has led to a situation in which current methods are not applicable to the amount of data available and researchers are forced to down-sample available data or to infer parameters from insufficient summary statistics. This research project will address the need for optimally designed coalescent modeling for inference from modern molecular data. The coalescent is a probability model on genealogies, that is, the trees which represent the ancestry of the sample. Coalescent models are used for inferring parameters of scientific relevance such as effective population size, migration patterns and selection. The research goals of this project are to expand the class of coalescent models and to design novel efficient statistical algorithms, allowing us to address many practical problems that advance science. Furthermore, the outcomes of the projects will foster the development of new statistical theory and tractable methods that contribute to biological solutions. This project also outlines an active plan for a broad range of educational and outreach activities that will broaden participation in statistical sciences and will enhance more inclusive atmosphere in science. The undergraduate and graduate students involved into the project will be offered a unique opportunity for interdisciplinary hands-on research training at the interface of statistical sciences and biology, allowing them to contribute to progress in evolutionary biology, molecular biology, population genetics, phylogenetics, cancer genomics, probabilistic modeling, statistical inference, and related fields. The PI will actively participate in multiple outreach activities such as the Stanford undergraduate summer research program, which will allow for recruiting more diverse pool of future data scientists and for fostering more inclusive climate in science. The research findings of the project will serve as foundation for new program in statistical genetics and will be integrated into undergraduate and graduate courses. Concretely, this project will expand the class of coalescent models and provide a suite of new algorithmic and statistical approaches by exploiting a metric notion of genealogies, lumpability of Markov chains and divide-and-conquer strategies. The specific aims include (1) develop coalescent models to incorporate various sampling schemes and biological processes such as dynamic population structures, recombination and strong selection; (2) develop a metric framework for coalescent theory and applications; (3) develop scalable strategies for Bayesian inference of evolutionary parameters and (4) implement, validate and analyze molecular sequences of infectious disease such as SARS-CoV-2, ancient and modern human DNA samples and cancer single cell variation. Furthermore, the project will actively contribute to broadening participation in statistical sciences at multiple fronts, from team-based interdisciplinary research training to community outreach.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1093/biomet/asad025
发表时间:
2023-06-23
期刊:
BIOMETRIKA
影响因子:
2.7
作者:
[Samyak,Rajanala, Palacios,Julia A.]
通讯作者:
Palacios,Julia A.
DOI:
10.1093/jrsssc/qlad098
发表时间:
2024-03-11
期刊:
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES C-APPLIED STATISTICS
影响因子:
1.6
作者:
[Zhang,Julie, Preising,Gabriel A., Palacios,Julia A.]
通讯作者:
Palacios,Julia A.
海外基金