课题基金 / 基金详情

Leveraging large language models and knowledge graphs on clinical, pathological, and sequencing data to inform precision cancer therapy

Leveraging large language models and knowledge graphs on clinical, pathological, and sequencing data to inform precision cancer therapy
利用临床、病理和测序数据的大型语言模型和知识图为精准癌症治疗提供信息
批准号:
10888730
负责人:
SELWYN M VICKERS
金额:
$30.0万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-01-20 至 2023-12-31

项目摘要

项目成果

SELWYN M VICKERS的其他基金

相关文献

中文摘要
翻译
项目总结: 精确医学和靶向治疗是癌症生物学中的新兴领域,旨在将 个人层面的临床、病理和基因组概况,为癌症患者量身定做治疗策略。 已经建立了几个精确的肿瘤学知识库,如OncoKB,我的癌症基因组,以 通过利用专家对生物和临床意义的管理,使临床决策民主化 使用公开可用的资源进行更改。这些知识库虽然非常强大,但也有其 限制,包括注释基因和改变的范围,以及确定准确的治疗方法 患者基因组和临床特征的特定组合。在这项提案中,我们计划开发新的 将整合(I)广泛的隐含癌症知识的计算方法 通过具有(Ii)明确结构的临床、病理和基因组的大型语言模型(LLM) 斯隆·凯特琳纪念癌症中心(MSKCC)癌症患者的知识 临床测序队列和AACR项目精灵队列。专家将进一步加强这一点 筛选,目的是预测基因组变化和临床或病理特征的组合 这可以与一种特定的癌症疗法相匹配。这项研究的目标是开发计算 模型从根本上以知识图和LLM为基础,以弥合临床和 癌症和癌症治疗的功能性风险因素,并提供信息和加强个性化治疗。 这项建议的第一个目标是开发一个知识图谱,MSK-CancerKG,基于患者特定的临床, 来自MSKCC临床的100,000多名患者的病理和基因组改变信息 测序队列和AACR精灵项目队列。该多关系知识图谱将集成一个 与每个患者相关的广泛的临床特征,从病理报告中提取特征 与患者来源的肿瘤样本相对应,以及基因组的全面特征 改变和牵连的基因。第二个目标将针对预先训练的大型 使用MSK的结构化、详细和更可靠的癌症特定知识的语言模型(LLM)- 巨蟹座。我们将仔细地将这些微调的模型与4个经过预先培训的最先进的模型进行比较 语言模型,最终得出一个优化的组合预测模型,创造了MSK-CancerLLM。这个 基准步骤将包括成功的临床、改变和治疗预测准确性 病人数据。该提案的第三个目标将是利用临床实践进一步微调MSK-CancerLLM 指导方针和反馈,以模拟癌症领域专家的输出。生成的模型将集成到 一个名为MSK-Assistant的AI聊天机器人,用于促进后端之间的无缝集成和交互 模型和前端聊天机器人界面。与ChatGPT应用程序一样,这将允许研究社区 关于癌症生物学以及个性化药物建议和治疗干预的询问。
英文摘要
Project Summary: Precision medicine and targeted therapy are emerging domains in cancer biology that aim to incorporate individual-level clinical, pathological and genomic profiles to tailor treatment strategies for cancer patients. Several precision oncology knowledge bases, like OncoKB, My Cancer Genome, have been established to democratize clinical decision-making by leveraging expert curation of biological and clinical significance of alterations using publicly available resources. These knowledge bases, while extremely powerful, have their limitations, including the scope of annotated genes and alterations, as well as identifying precise therapies for specific combinations of a patient's genomic and clinical profiles. In this proposal, we plan to develop new computational methodologies that will integrate (i) the broad range of implicit cancer knowledge accrued by Large Language Models (LLMs) with (ii) the explicit structured clinical, pathological, and genomic knowledge derived from cancer patients in the Memorial Sloan Kettering Cancer Center’s (MSKCC) Clinical Sequencing cohort and AACR Project GENIE cohort. This will further be reinforced by expert curation, with the aim to predict combinations of genomic alterations and clinical or pathological profiles that can be matched to a specific cancer therapy. The goal of this research is to develop computational models fundamentally anchored around knowledge graphs and LLMs to bridge the gap between clinical and functional risk factors of cancer and cancer therapeutics, and to inform and enhance personalized therapies. The first aim of this proposal is to develop a knowledge graph, MSK-CancerKG, based on patient-specific clinical, pathological, and genomic alteration information from more than 100,000 patients from the MSKCC Clinical Sequencing Cohort and the AACR GENIE Project cohort. This multi-relational knowledge graph will integrate a wide spectrum of clinical features associated with each patient, abstracted features from pathological reports corresponding to the patient-derived tumor samples, along with comprehensive characterization of genomic alterations and the implicated genes. The second aim will be geared towards the fine-tuning of pre-trained Large Language Models (LLMs) using the structured, detailed and more reliable cancer-specific knowledge from MSK- CancerKG. We will meticulously benchmark these fine-tuned models against 4 state-of-the art pre-trained language models, ultimately deriving an optimized combined predictive model, coined MSK-CancerLLM. The benchmarking step will include successful clinical, alteration and treatment prediction accuracy on held-out patient data. The third aim of the proposal will be to further fine-tune MSK-CancerLLM using clinical practice guidelines and feedback to model output from cancer domain experts. The resulting model will be integrated into an AI chatbot, called MSK-Assistant, to facilitate seamless integration and interaction between the backend model and a frontend chatbot interface. Like the ChatGPT application, this will allow the research community to query about cancer biology and personalized drug recommendations and therapeutic interventions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
UAB/TU FIRST Administrative Core
UAB/TU FIRST Administrative Core
Clinical Managment and Trials Core and Advocacy Sub-Core
Research Training/Education Core
  • 批准号:
    7771813
  • 项目类别:
  • 资助金额:
    $19.0万
  • 财政年份:
    2009
  • 负责人:
    SELWYN M VICKERS
  • 依托单位: