课题基金 / 基金详情

Implicit generative modeling for computational genomics

Implicit generative modeling for computational genomics
计算基因组学的隐式生成模型
批准号:
RGPIN-2020-06770
负责人:
Delong, Andrew
金额:
$2.91万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31

项目摘要

项目成果

Delong, Andrew的其他基金

相似基金

相关文献

中文摘要
翻译
基因组学是研究细胞中的DNA序列及其如何驱动功能或功能障碍的学科。基因组学数据还揭示了细胞内的信息流。通过比较在不同条件(突变或药物治疗)下捕获的基因组数据,我们可以观察它们如何影响细胞功能。机器学习可以做的不仅仅是观察:它可以直接从数据中学习DNA驱动器运行所依据的复杂“规则”。一旦学习,这些规则就可以预测哪些突变可能导致基因停止功能并导致疾病。基因组学数据越来越多地可以由自动化实验室产生,包括远程云实验室,它们提供了比传统实验室工作更好的准确性和一致性。这种信息丰富的数据和自动化实验的结合终于为“自动驾驶”实验室铺平了道路。在自我驱动的实验室里,实验是由机器学习系统设计的,以系统地询问感兴趣的结果,例如人类细胞中的序列-功能关系。数据和模型在全天候反馈循环中得到改进。拟议的研究旨在朝着这一愿景取得可处理的进展。这一愿景很重要,因为它代表着药物发现的机会。今天,90%的候选药物在经过多年的开发和数千万美元的花费后,在临床阶段失败了。对细胞如何工作的更好的预测模型将有助于提前识别和设计针对遗传疾病的化合物,最终成功率更高。拟议的研究集中在四个目标上。首先是让机器学习科学家更容易参与推进这一愿景。基因组学数据是出了名的难以理解和使用,所以我们的目标是建立一个代码库和数据存储库,使机器学习社区能够为基因组学研究做出贡献。第二个目标是提高解释基因组学数据的准确性,特别是来自RNA测序或RNA序列的数据。Rna-seq捕获样本中的rna分子的“快照”,是无数询问细胞方案中的倒数第二步。但是,快照是零碎的。需要算法来推断(重建)样本中实际存在的RNA分子。这项拟议的研究将探索一种新的、更准确的基于深度学习的推理方法。第三个目标是通过新的深度学习技术建立更准确的RNA在细胞中处理的模型。RNA加工模型对于预测突变和潜在疗法的效果很重要。(这些模型也以RNA-seq数据为基础,并可能受益于第二个目标的进展。)第四个也是最后一个目标是创建能够自动设计实验的算法,生成最有价值的数据,以此来建立RNA生物学模型。
英文摘要
Genomics is the study of DNA sequences in our cells and how they drive function or dysfunction. Genomics data also reveals the flow of information within cells. By comparing genomics data captured under different conditions (mutations or drug treatments) we can observe how they affect cell function. Machine learning can do more than merely observe: it can learn the complex "rules" by which DNA drives function, directly from data. Once learned, these rules can predict which mutations are likely to cause a gene to stop functioning and cause disease. Genomics data can increasingly be generated by automated laboratories, including remote cloud laboratories, which offer better accuracy and consistency than traditional laboratory work. This combination of information-rich data and automated experimentation is finally setting the stage for "self-driven" laboratories. In a self-driven laboratory, experiments are designed by machine learning systems to systematically interrogate an outcome of interest, such as the sequence-function relationships in human cells. Data and models are improved in a 24/7 feedback loop. The proposed research aims to make tractable progress towards this vision. This vision is important because it represents an opportunity for drug discovery. Today, 90% of all drug candidates fail at the clinical stage, after years of development and tens or hundreds of millions of dollars have been spent. Better predictive models of how cells work will help to identify and design compounds for genetic disorders up front, with a higher final success rate. The proposed research focuses on four objectives. The first is to make it easier for machine learning scientists to participate in advancing this vision. Genomics data is notoriously difficult to understand and to work with, so the goal is to build a code library and data repository that enables the machine learning community to contribute to genomics research. The second objective is to improve the accuracy with which genomics data is interpreted, specifically data from RNA sequencing, or RNA-seq. RNA-seq captures a "snapshot" of the RNA molecules in a sample, and is the penultimate step in countless protocols for interrogating cells. However, the snapshot is fragmented. Algorithms are needed to infer (to reconstruct) what RNA molecules were actually present in the sample. The proposed research will investigate a new, more accurate approach to inference based on deep learning. The third objective is to build more accurate models of how RNA is processed in cells by new deep learning techniques. Models of RNA processing are important for predicting the effects of mutations and of potential therapies. (Such models are also based on RNA-seq data, and may benefit from advances in the second objective.) The fourth and final objective is to create algorithms that can automatically design experiments, generating the most valuable data possible from which to build models of RNA biology.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Implicit generative modeling for computational genomics
  • 批准号:
    RGPIN-2020-06770
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.91万
  • 财政年份:
    2022
  • 负责人:
    Delong, Andrew
  • 依托单位:
Implicit generative modeling for computational genomics
  • 批准号:
    RGPIN-2020-06770
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.91万
  • 财政年份:
    2020
  • 负责人:
    Delong, Andrew
  • 依托单位:
Implicit generative modeling for computational genomics
  • 批准号:
    DGECR-2020-00323
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2020
  • 负责人:
    Delong, Andrew
  • 依托单位:
Learning and Inference for Hierarchical Models in Computer Vision
  • 批准号:
    421338-2012
  • 项目类别:
    Postdoctoral Fellowships
  • 资助金额:
    $2.91万
  • 财政年份:
    2013
  • 负责人:
    Delong, Andrew
  • 依托单位:
海外基金