课题基金 / 基金详情

CAREER: Concise descriptors of genomic data facilitate mechanistic inference

CAREER: Concise descriptors of genomic data facilitate mechanistic inference
职业:基因组数据的简明描述符有助于机械推理
批准号:
2238125
负责人:
Maria Chikina
金额:
$75.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-01 至 2028-07-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
生物系统不仅仅是它们各部分的总和。现代高通量技术使科学家能够通过在单个样本中测量数千到数百万个分子测量来整体研究生物系统。这一新的科学范式加速了我们对复杂生物过程的理解,如胚胎发育和癌症。然而,需要复杂的分析框架才能将大型数据集转化为具体的生物学见解。该项目旨在建立通用和灵活的算法,可以将大量数据减少为更小的具有生物意义的表示。这些表示形式足够简洁,便于非专家操作,但也足够丰富,可以同时支持最初的预期分析和未来的数据重用。该项目还将开发教材和动手活动,以培养不同教育水平和技术背景的下一代生物数据科学家。该项目将开发一个高度可定制的贝叶斯框架,通过使用灵活的混合分布先验,将许多可解释降维的方法结合到一个统一的框架中。此外,利用丰富的生物学先验知识来源,所提出的方法将能够量化和按名称识别嵌入高维数据中的特定生物过程。在一个单独但协同的目标中,该项目将采用最先进的分割算法,从多样本、基本分辨率的DNA甲基化数据集中自动提取基因组特征,加快下游分析和解释。这项工作的结果可以在http://chikinalab.org/.This上找到,该奖项反映了国家科学基金会的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Biological systems are more than the sum of their parts. Modern high throughput technologies enable scientists to study biological systems holistically by measuring thousands to millions of molecular measurements in a single sample. This new scientific paradigm has accelerated our understanding of complex biological processes such as embryonic development and cancer. However, sophisticated analytical frameworks are needed to turn large datasets into concrete biological insights. This project aims to build general and flexible algorithms that can reduce large collections of data into smaller biologically meaningful representations. These representations are concise enough to be easily manipulated by non-experts yet rich enough to support both the originally intended analysis and future data reuse. The project will also develop teaching materials and hand-on activities to educate the next generation of biological data scientists across diverse education levels and technical backgrounds.This project will develop a highly customizable Bayesian framework that combines many approaches to interpretable dimensionality reduction into a single unifying framework by using flexible mixture distribution priors. Additionally, drawing on rich sources of biological prior knowledge the proposed method will be capable of both quantifying and identifying by name specific biological processes embedded in high dimensional data. In a separate but synergistic goal the project will adapt state-of-the-art segmentation algorithms to automatically extract genomic features from multi-sample, base-resolution DNA methylation datasets accelerating downstream analysis and interpretation. The results of this work can be found at http://chikinalab.org/.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金