TRTech-PGR: Connecting sequences to functions within and between species through computational modeling and experimental studies
TRTech-PGR: Connecting sequences to functions within and between species through computational modeling and experimental studies
批准号:
2107215
负责人:
Shin-Han Shiu
金额:
$140.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-12-01 至 2024-11-30
中文摘要
我们所知道的生命如果没有植物就不可能存在。它们是食物、氧气、木材、纤维和药物的来源。因此,提高植物性状,如产量、营养品质和抗逆性,对植物产品的可持续生产至关重要。我们改良植物的关键在于彻底了解植物DNA是如何控制性状的。例如,玉米DNA包含约20亿个字母,这些字母的不同组合影响不同的植物性状。但我们对哪些字母起作用以及它们如何控制性格的了解有限。当我们确实对DNA和性状之间的联系有了很好的理解时,这种理解仅限于少数几种模式植物,因为它们相对容易研究。因此,为了更全面地了解植物是如何工作的,我们将使用基于人工智能的方法将DNA序列与它们控制的性状联系起来,机器学习中使用计算机从广泛的生物数据中发现隐藏的模式。此外,我们将应用迁移学习将知识从一种植物物种转化为另一种植物物种,这样我们就可以将我们对模式植物的了解转移到其他物种。该项目的结果将是计算机程序,可以预测DNA序列和特征之间的联系,并在物种之间传递信息。利用这些程序,科学家们可以更好地了解植物是如何工作的,这些知识最终可以用来创造更多产、更有弹性的植物。组学数据的快速增长带来了改变植物科学的发现。然而,随着越来越多的基因组可用,将序列与其全球功能连接起来仍然具有挑战性。因此,我们的第一个目标是建立和验证可以预测序列函数的计算模型。第二个项目目标是开发和应用迁移学习来解决跨物种和环境的序列-功能问题。为了实现第一个目标,来自拟南芥、玉米、水稻和番茄四个模式物种的现有多组学和表型数据将与机器学习相结合,以解决两个序列到功能的问题:预测生物过程功能,如酶或信号通路成员,以及生理和形态表型。这些预测模型将使用模型解释方法进行剖析,通过理解模型工作的原因和方式来提供机制见解。为了实现我们的第二个目标,使用来自目标模型物种的相同数据并解决相同的焦点问题,迁移学习策略将被开发和优化,以评估如何在物种和环境之间最好地转移知识。我们将重点关注的四种模型有相对丰富的实验数据,通过持有不同数量和类型的数据,可以重新创建和评估各种“数据贫乏”的情景。对于这两个项目目标,预测将通过独立于建模数据和为本项目进行的遗传实验的新数据的实验数据进行验证。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Life as we know it would be impossible without plants. They are a source of food, oxygen, timber, fiber, and medicine. Therefore, improving plant traits, such as yield, nutritional quality, and resilience, is crucial for sustainable production of plant products. Key to our ability to improve plants is a thorough understanding of how plant DNA controls traits. For example, corn DNA contains ~2 billion letters, and different sets of these letters affect different plant traits. But we have limited knowledge about which letters matter and how they control traits. When we do have a good understanding of the connection between DNA and traits, such understanding is limited to a handful of model plants chosen for their relative ease of study. Thus, to have more complete knowledge of how plants work, we will connect DNA sequences with traits they control using an Artificial Intelligence-based approach, machine learning where computers are used to uncover hidden patterns from a wide range of biological data. In addition, we will apply transfer learning to translate knowledge from one plant species to another so we can later transfer what we know about model plants to other species. The outcome of the project will be computer programs that can predict the connections between DNA sequence and traits and transfer information across species. Using these programs, scientists can better understand how plants work and this knowledge can ultimately be used to create more productive and resilient plants.The rapid growth in omics data has led to discoveries transforming plant science. However, as more genomes become available, connecting sequences to their functions globally remains challenging. Thus, our first goal is to build and validate computational models that can predict sequence functions. The second project goal is to develop and apply transfer learning to address sequence-to-function problems across species and environments. To achieve the first goal, existing multi-omics and phenotype data from four model species–Arabidopsis, maize, rice, and tomato—will be integrated with machine learning to address two sequence-to-function problems: predictions of biological process functions such as enzyme or signaling pathway membership, and physiological and morphological phenotypes. These prediction models will be dissected using model interpretation methods to provide mechanistic insights through understanding why and how the models work. To achieve our second goal, using the same data from target model species and addressing the same focal problems, transfer learning strategies will be developed and optimized to assess how knowledge can be best transferred across species and environments. There is relatively abundant experimental data available for the four models we will focus on, and by holding out different amounts and types of data, a wide range of “data-poor” scenarios can be recreated and evaluated. For both project goals, the predictions will be validated with holdout experimental data independent from data used for modeling and new data from genetic experiments conducted for this project.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(16)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Plant science corpus
植物科学语料库
DOI:
10.5281/zenodo.10022686
发表时间:
2023
期刊:
Zenodo
影响因子:
--
作者:
[Shiu, Shin-Han]
通讯作者:
Shiu, Shin-Han
DOI:
10.1016/j.algal.2022.102709
发表时间:
2022-04-21
期刊:
ALGAL RESEARCH-BIOMASS BIOFUELS AND BIOPRODUCTS
影响因子:
5.1
作者:
[Lucker,Ben F., Temple,Joshua A., Kramer,David M.]
通讯作者:
Kramer,David M.
Plant Science Knowledge Graph Corpus: a gold standard entity and relation corpus for the molecular plant sciences
植物科学知识图谱语料库:分子植物科学的黄金标准实体和关系语料库
DOI:
10.1093/insilicoplants/diad021
发表时间:
2023
期刊:
in silico Plants
影响因子:
3.1
作者:
[Lotreck, Serena, Segura Abá, Kenia, Lehti-Shiu, Melissa D., Seeger, Abigail, Brown, Brianna N. I., Ranaweera, Thilanka, Schumacher, Ally, Ghassemi, Mohammad, Shiu, Shin-Han, Marshall-Colon, ed., Amy]
通讯作者:
Marshall-Colon, ed., Amy
DOI:
10.1093/biosci/biad015
发表时间:
2023-04-29
期刊:
BIOSCIENCE
影响因子:
10.1
作者:
[Cuddington,Kim, Abbott,Karen C., White,Easton R.]
通讯作者:
White,Easton R.
Supervised capacity preserving mapping: a clustering guided visualization method for scRNA-seq data.
保存映射的监督能力:用于SCRNA-SEQ数据的聚类指导性可视化方法。
DOI:
10.1093/bioinformatics/btac131
发表时间:
2022-04-28
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[]
通讯作者:
共 7 条
RESEARCH-PGR: Combining machine learning and experimental analysis to define trichome and root-specific gene regulatory networks in cultivated tomato and related Solanaceae species
-
批准号:2218206
-
项目类别:Standard Grant
-
资助金额:$180.0万
-
财政年份:2023
-
负责人:Shin-Han Shiu
-
依托单位:
Collaborative Research: Assessing the connections between genetic interactions, environments, and phenotypes in Arabidopsis thaliana
-
批准号:2210431
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2022
-
负责人:Shin-Han Shiu
-
依托单位:
NRT-HDR: Intersecting computational and data science to address grand challenges in plant biology
-
批准号:1828149
-
项目类别:Standard Grant
-
资助金额:$300.0万
-
财政年份:2018
-
负责人:Shin-Han Shiu
-
依托单位:
Collaborative Research: Fitness effects of loss-of-function mutations in duplicate genes
-
批准号:1655386
-
项目类别:Standard Grant
-
资助金额:$59.4万
-
财政年份:2017
-
负责人:Shin-Han Shiu
-
依托单位:
Computational and Experimental Studies of Plastid Functional Networks
-
批准号:1119778
-
项目类别:Continuing Grant
-
资助金额:$122.22万
-
财政年份:2011
-
负责人:Shin-Han Shiu
-
依托单位:
Experimental Characterization of Novel Coding Small ORFs in the Arabidopsis thaliana Genome
-
批准号:0749634
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2008
-
负责人:Shin-Han Shiu
-
依托单位:
国内基金
海外基金
登录
查看更多内容
TET2去甲基化上调CAV1表达介导PGR泛素化降解在妊娠期显性糖尿病并发子痫前期蜕膜化障碍中的作用及干预研究
-
批准号:JCZRLH202600862
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
E3连接酶RNF213导致PGR缺陷在子宫内膜蜕膜化中的作用机制研究
-
批准号:--
-
项目类别:地区科学基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:林忠
-
依托单位:
孕激素通过 PGR/RUNX 调控胎盘 ASPROSIN 转录介
导妊娠期糖尿病
-
批准号:2024JJ5350
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:洪涛
-
依托单位:
通过构建Pgr-Cas9工具小鼠研究Hippo通路效应因子Yap1/Wwtr1在蜕膜化过程中的作用
-
批准号:32370913
-
项目类别:面上项目
-
资助金额:50万元
-
批准年份:2023
-
负责人:刘极龙
-
依托单位:
海洋硅藻PGR5/PGRL1蛋白感知和适应波动光的作用机制研究
-
批准号:42276146
-
项目类别:面上项目
-
资助金额:56万元
-
批准年份:2022
-
负责人:王广策
-
依托单位:
KLF12通过调控PGR和GDF10的表达抑制孕激素诱导子宫内膜癌细胞分化的机制研究
-
批准号:--
-
项目类别:面上项目
-
资助金额:55万元
-
批准年份:2021
-
负责人:周怀君
-
依托单位:
HBP1调节PGR转录活性在胚胎植入及妊娠维持中的作用机制
-
批准号:82160296
-
项目类别:地区科学基金项目
-
资助金额:34.00万元
-
批准年份:2021
-
负责人:黄品秀
-
依托单位:
靶向PGR阳性乳腺癌的多功能钌配合物合成及其抗肿瘤机制研究
-
批准号:21501074
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2015
-
负责人:吕高超
-
依托单位: