课题基金 / 基金详情

CAREER: Practical algorithms and high dimensional statistical methods for multimodal haplotype modelling

CAREER: Practical algorithms and high dimensional statistical methods for multimodal haplotype modelling
职业:多模态单倍型建模的实用算法和高维统计方法
批准号:
2239870
负责人:
Derek Aguiar
金额:
$54.83万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-15 至 2028-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
人类细胞已经产生了大量和多样化的数据集,目的是解释细胞差异如何影响人与人之间观察到的特征差异。例如,人与人之间遗传差异的数学模型可以用来解释为什么有些人容易患上某种疾病。然而,大多数数学模型对遗传差异如何相互作用影响观察到的特征做出了过于简单化的假设。该项目通过创新新的稳健的数学模型来解决计算生物学和应用机器学习中的主要挑战,这些模型几乎不做假设,并且使用高效的训练算法来利用海量和复杂的细胞数据。具体地说,该项目考虑:(A)通过整合不同类型的数据、机器学习和算法技术来计算遗传差异序列的方法;(B)表征人与人之间遗传相似性的数学模型;以及(C)大规模数据集的有效算法。该项目的结果包括广泛适用于对海量和多样化的序列数据进行聚类的新方法,尤其有助于研究人员试图了解遗传差异如何影响疾病和其他特征。此外,该研究通过开发互动学习模块和网络资源来支持数学和科学高中和大学社区。本项目在两个研究方向上发展了多峰变异序列(即多染色体单倍型)的统计和算法基础。第一个方向介绍了多体单倍型数据结构,并发展了新的贝叶斯非参数模型和快速推理算法,用于从异质和高维生物分子数据中聚类多体单倍型。通过在数据空间(贝叶斯核集)、模型空间(深度近似)和算法空间(变分近似)中操作的新颖而高效的推理算法来实现计算可处理性。第二个方向发展了第一个模型,该模型将单倍型组装的组合域与概率单倍型阶段化域统一起来,以推断潜在的单倍型。研究人员将通过将有向和无向图形建模技术与高效的基于粒子的推理算法相结合来实现这一统一目标。这些研究任务的完成将产生新的方法来开发高维贝叶斯非参数模型的深度近似、多峰序贯聚类模型以及加速高维统计模型的训练。此外,这项研究解决了(A)单倍型组装和单倍型阶段化统一的长期悬而未决的问题;以及(B)关联性研究中遗漏遗传性的潜在来源:相依遗传和单倍型-表观遗传相互作用。与大学和地区高中社区的合作将把研究成果转化为教育模块和资源,以激励、吸引和留住计算机科学学生和教师。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Massive and diverse datasets have been generated from human cells with the goal of explaining the many ways cellular differences affect the observed differences in traits between people. Mathematical models of the genetic differences between people can be used to explain, for example, why some individuals are predisposed to developing a particular disease. However, most mathematical models make overly simplistic assumptions about how genetic differences interact to influence an observed trait. This project addresses major challenges in computational biology and applied machine learning by innovating new robust mathematical models that make few assumptions and efficient training algorithms to leverage massive and complex cellular data. Specifically, the project considers: (a) methods for computing sequences of genetic differences by integrating different types of data, machine learning, and algorithmic techniques; (b) mathematical models for characterizing the genetic similarity between people; and (c) efficient algorithms that scale to large datasets. The results of this project include new methods that are broadly applicable to clustering massive and diverse sequential data, and specifically helpful for researchers trying to understand how genetic differences affect disease and other traits. Furthermore, the research supports the math and science high school and university communities by developing interactive learning modules and networking resources.This project develops the statistical and algorithmic foundations for sequences of multimodal variation (i.e., multiomic haplotypes) in two research directions. The first direction introduces the multiomic haplotype data structure and develops new Bayesian nonparametric models and fast inference algorithms for clustering multiomic haplotypes from heterogeneous and high dimensional biomolecular data. Computational tractability is achieved through novel and efficient inference algorithms that operate in data-space (Bayesian coresets), model-space (deep approximations), and algorithm-space (variational approximations). The second direction develops the first model that unifies the combinatorial domain of haplotype assembly with the probabilistic haplotype phasing domain to infer latent haplotypes. The investigator will accomplish this unification goal by combining directed and undirected graphical modeling techniques with efficient particle-based inference algorithms. The completion of these research tasks will result in new methods for developing deep approximations for high dimensional Bayesian nonparametric models, models for multimodal sequential clustering, and methods to accelerate the training of high dimensional statistical models. Additionally, the research addresses (a) the longstanding open problem of haplotype assembly and haplotype phasing unification; and (b) potential sources of missing heritability in association studies: phase-dependent genetic and haplotype-epigenetic interactions. Partnerships with the university and regional high school communities will translate the research findings into educational modules and resources to motivate, engage, and retain computer science students and teachers.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金