课题基金 / 基金详情

CRII: III: Self-Supervised Graph Neural Network Meta-Learning for Cancer Multi-Omics and Driver Discovery

CRII: III: Self-Supervised Graph Neural Network Meta-Learning for Cancer Multi-Omics and Driver Discovery
CRII:III:用于癌症多组学和驱动发现的自监督图神经网络元学习
批准号:
2245805
负责人:
Tianle Ma
金额:
$15.69万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-05-01 至 2025-04-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
癌症是全球第二大死亡原因,每年造成800多万人死亡。癌症是由驱动突变引起的,驱动突变是基因DNA序列的变化,导致异常细胞的发展,这些细胞不受控制地分裂,并可以渗透和破坏正常的身体组织。癌症研究的一个主要目标是发现所有癌症类型的驱动突变。然而,癌症驱动基因的发现非常具有挑战性,因为每个人的癌症都有独特的突变组合,其中只有少数是驱动基因突变,而绝大多数是乘客突变,不会促进异常细胞生长。现有的癌症驱动因子发现方法可以检测许多肿瘤中常见的驱动因子,但无法检测仅在极少数肿瘤病例中出现的罕见驱动因子。该项目将应用最先进的机器学习方法来识别剩余的难以捉摸的罕见驱动程序。识别的驱动突变可能为计划治疗提供关键信息,以阻止癌细胞生长,包括靶向特定驱动突变的药物。从该项目中获得的知识和框架将有助于通过提供更好地管理癌症治疗的能力来挽救许多生命。将深度神经网络模型等先进机器学习方法应用于癌症驱动因素发现的一个主要障碍是缺乏适合监督学习的大规模高质量标记训练数据。为了应对这一挑战,该项目将结合联合收割机图神经网络(GNN),自我监督学习和Meta学习技术,将生物领域知识与大规模异构数据(在该社区中具体称为多组学数据)整合,用于癌症驱动因素发现。首先,基于表示生物领域知识的统一知识图构建GNN模型。将领域知识作为归纳偏差的一种形式引入模型中,将有助于用更少的标记数据有效地训练模型。其次,自监督学习将用于在多组学数据上预训练GNN模型,也减少了对标记数据的需求。同时,GNN模型的学习节点和边嵌入可以被视为高级可转移特征,消除异质性和噪声,同时促进跨数据集集成和Meta学习。第三,Meta学习将应用于包括数十种癌症类型的泛癌症数据集,以提高模型的可推广性并检测癌症类型的新驱动因素。该项目将进一步改进现有的对数十种癌症类型的泛癌症综合分析的结果,并可能通过将异构数据和知识源与机器学习相结合,为解决其他困难的生物学问题带来可重复的过程。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Cancer is the second leading cause of death worldwide, killing more than 8 million people every year. Cancer is caused by driver mutations, which are changes in the DNA sequence of genes that lead to the development of abnormal cells that divide uncontrollably and can infiltrate and destroy normal body tissue. One main goal of cancer research has been the discovery of all the driver mutations across cancer types. However, cancer driver discovery is very challenging as each person’s cancer has a unique combination of mutations, among which only a few are driver mutations while the vast majority are passenger mutations that do not promote abnormal cell growth. Existing cancer driver discovery methods can detect drivers that are common among many tumors but fail to detect rare drivers that only occur in very few tumor cases. This project will apply state-of-the-art machine learning approaches to identify the remaining elusive rare drivers. The identified driver mutations may provide critical information for planning treatment to stop cancer cells from growing, including drugs that target a specific driver mutation. The acquired knowledge and framework from this project will contribute to saving numerous lives by providing the capability to better manage cancer treatments. One major barrier for applying advanced machine learning approaches, such as deep neural network models, to cancer driver discovery is the lack of large-scale high-quality labeled training data suitable for supervised learning. To address this challenge, this project will combine Graph Neural Network (GNN), Self-Supervised Learning, and Meta Learning techniques to integrate biological domain knowledge with large-scale heterogeneous data (specifically known as multi-omics data in this community) for cancer driver discovery. First, a GNN model will be constructed based on a unified knowledge graph representing biological domain knowledge. Incorporating domain knowledge into the model as a form of inductive bias will help train the model effectively with less labeled data. Second, self-supervised learning will be employed to pre-train the GNN model on multi-omics data, also reducing the need for labeled data. Meanwhile, the learned node and edge embeddings for the GNN model can be treated as high-level transferrable features, removing heterogeneity and noise while facilitating cross-dataset integration and meta learning. Third, meta learning will be applied to a pan-cancer dataset comprising dozens of cancer types to increase model generalizability and detect novel drivers across cancer types. The project will further improve the results of existing pan-cancer integrative analysis of dozens of cancer types and may lead to a repeatable process for tackling other difficult biological problems through integrating heterogeneous data and knowledge sources with machine learning.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
  • 批准号:
    JCZRLH202600780
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
  • 批准号:
    2026JJ82690
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张卓
  • 依托单位:
基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
  • 批准号:
    2026JJ30130
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张二军
  • 依托单位: