CAREER: Scalable algorithms for regularized and non-linear genetic models of gene expression
CAREER: Scalable algorithms for regularized and non-linear genetic models of gene expression
批准号:
2336469
负责人:
Tiffany Amariuta-Bartell
金额:
$60.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-03-01 至 2029-02-28
中文摘要
DNA突变对基因的工作方式有深远的影响,但人们仍然不太清楚哪些突变会影响哪些基因。目前,由于在分析基因组数据方面的挑战,我们的知识有限,例如由于欧洲研究参与者的过度代表性而产生的偏见,以及不能充分捕捉数据的简单统计模型。这个项目跨越了三个主要的科学目标和两个教育目标,一个是创新的统计模型将DNA突变映射到他们的目标基因,另一个是同时培养科学培训和多样性。首先,研究人员将提高将突变映射到目标基因的保真度,以定位那些没有得到很好研究的个体群体,例如少数族裔群体。其次,研究人员将开发一种新的方法,通过考虑基因如何在全基因组网络中相互作用,将突变与基因联系起来,这表明许多未确定特征的突变具有功能效应。第三,研究人员将使用反映单细胞基因组分析数据自然分布的可扩展模型来表征突变发挥影响的特定细胞。这项研究通过引入将突变与目标基因联系起来的新的、强大的统计模型,推动了生物信息学和人类遗传学领域的发展。该项目还加强了生物医学发现的公平性和多样性,同时提高了研究环境中的多样性。对于后者,调查人员为来自资源不足社区的高中生启动了一个为期数周的校园研究项目,并为本科生和研究生开设了遗传学培训课程,提供业界和学术界梦寐以求的定量跨学科技能。这一奖项将产生广泛的数据集、开放源码统计模型和基因组学工具、高影响力的出版物和课程材料,从而吸引和推动科学界参与和推动相关研究。这个项目的重点是开发新的遗传模型,以了解特定的遗传变异如何影响基因表达。这些模型克服了目前在表征遗传变异功能方面的局限性,遗传变异的后续目标往往是在调节人类表型(如身高和癌症风险)中牵涉到靶基因。现有算法的挑战包括由于样本大小有限(特别是对于研究不足的少数群体)的统计问题,限制从全基因组分析获得的知识的多个假设负担,以及模型错误指定,特别是对于日益流行的新数据类型,如单细胞基因组学。调查人员通过三个主要目标应对这些挑战。首先,研究人员通过联合对全球不同数据集的遗传关联进行建模,将基因变异与研究不足的少数群体中基因表达的变化联系起来。其次,研究人员开发了一种全面的方法,利用基因调控网络的先验知识和先进的机器学习算法,将全基因组范围的遗传变异映射到基因表达的变化,以减轻多次测试的负担。第三,研究人员设计了一种新的统计模型,以高分辨率表征基因表达调控的细胞类型特异性;该模型利用单细胞数据的自然分布,解决了模型对最先进方法的错误说明,并通过对捐赠者之间数百万个单细胞测量进行建模来减少测量噪声。该奖项支持开发以基因变异功能为特征的开源基因组学软件和数据存储库,同时也为资源不足的高中生和有动力的本科生和研究生创造教育和培训机会。共生研究和教育的关系交织在一起,预计将增强研究环境的多样性,以及研究队列的多样性。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
DNA mutations have a profound effect on how genes work, but it’s still not well understood which mutations affect which genes. Currently, our knowledge is limited due to challenges in analyzing genomics data, such as bias arising from an overrepresentation of European study participants and simplistic statistical models that do not sufficiently capture the data. This project overcomes these challenges across three main scientific goals, in which innovative statistical models map DNA mutations to their target genes, and two educational goals, in which scientific training and diversity are simultaneously cultivated. First, the investigators will improve the fidelity of mapping mutations to target genes for groups of individuals that are not well-studied, such as minority populations. Second, the investigators will develop a new method to connect mutations to genes by considering how genes interact with each other in genome-wide networks, suggesting functional effects for many uncharacterized mutations. Third, the investigators will characterize the specific cells in which mutations exert their effects using scalable models that reflect the natural distribution of data from single cell genomic assays. This research advances the fields of bioinformatics and human genetics by introducing new, robust statistical models that link mutations to their target genes. This project also enhances equity and diversity in biomedical discoveries, while simultaneously enhancing diversity within research environments. Toward the latter, the investigators initiate a multi-week on-campus research program for high school students from under-resourced communities, as well as genetics training courses for undergraduate and graduate students, supplying quantitative interdisciplinary skills coveted by industry and academia alike. This award will generate extensive datasets, open-source statistical models and genomics tools, high-impact publications, and course materials, thereby engaging and fueling the scientific community to partake and propel related research. This project focuses on developing new genetic models to understand how specific genetic variations influence gene expression. These models overcome current limitations in characterizing the function of genetic variation, which often has the subsequent goal of implicating target genes in the regulation of human phenotypes such as height and cancer risk. Challenges of existing algorithms include statistical issues due to finite sample sizes (especially for understudied minority populations), multiple hypothesis burdens restricting the knowledge gained from genome-wide analysis, and model misspecification especially for new datatypes of growing popularity, such as single cell genomics. The investigators address these challenges across three main objectives. First, the investigators link genetic variation to changes in gene expression in understudied minority populations, by jointly modeling genetic associations across globally diverse datasets. Second, the investigators develop a comprehensive approach to map genome-wide genetic variants to changes in gene expression using a priori knowledge of gene regulatory networks and advanced machine learning algorithms to reduce the burden of multiple testing. Third, the investigators design a new statistical model to characterize the cell-type-specificity of gene expression regulation at high resolution; this model leverages the natural distribution of single cell data, resolving model misspecification of state-of-the-art methods and reduces measurement noise by modeling millions of single cell measurements across donors. This award supports the generation of open-source genomics software and data repositories characterizing the function of genetic variants, while also creating educational and training opportunities for under-resourced high school students and motivated undergraduate and graduate students. The symbiotic research and educational intertwine in a relationship that is expected to enhance both the diversity in research environments, as well as the diversity in research cohorts.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位: