课题基金 / 基金详情

Applying Large Language Models to Accelerate Abstraction of Cancer Pathology Reports for Cancer Registry (LLMs for Unstructured Data Extraction)

Applying Large Language Models to Accelerate Abstraction of Cancer Pathology Reports for Cancer Registry (LLMs for Unstructured Data Extraction)
应用大型语言模型加速癌症登记的癌症病理报告的抽象(非结构化数据提取法学硕士)
批准号:
10890243
负责人:
John L. Cleveland
金额:
$30.0万
依托单位国家:
美国
项目类别:
财政年份:
1998
资助国家:
美国
项目状态:
未结题
起止时间:
1998-02-18 至 2027-01-31

项目摘要

项目成果

John L. Cleveland的其他基金

相似基金

相关文献

中文摘要
翻译
病理报告,包含组织样本和病变的关键信息,在
英文摘要
Pathology reports, containing critical information on tissue samples and lesions, play a significant role in determining cancer treatment selection, prognosis, risk stratification, and clinical trial screening. Yet, manually extracting tumor characteristics from these unstructured or semi-structured reports is a complex, laborious process. Recent advances in Natural Language Processing (NLP) via deep learning methodologies show promising potential. Though Bidirectional Encoder Representations from Transformers (BERT) has achieved notable results in various NLP tasks, its application in pathology is constrained due to the limited allowable input length. Our recent study addressed this by transfer learning a BERT-based model on increasingly complex knowledge sources including Wikipedia, PubMed, MIMIC-III, and Moffitt institutional pathology reports. This language model was further fine-tuned to identify site, histology, and associated ICD-O-3 codes from pathology reports. Despite the promising preliminary results, our pilot work focuses on extractive question-anwsering of single primary solid tumor diagnosis, overlooking rich terminology and variation of the pathology language. Our long-term goal is to employ Large Language Models (LLMs) to extract information from all types of clinical notes, assisting institutional certified tumor registrars in data abstraction for the Cancer Registry. In this work, we specifically focus on pathology reports, and proposes to train LLMs on 349,544 institutional pathology reports to identify five key cancer data elements: primary site, histology, stage, grade, and laterality. The study will focus on common (breast) and rare (gastric) cancers. We will leverage existing LLMs pretrained on large public corpora, retrain them on institutional pathology reports, and finally fine-tune them to predict specific cancer data elements. We pursue two specific aims. Aim 1: predict breast cancer data elements by abstractive question-aswering using the existing cabernet architecture (Aim 1a), and by a prompt-based finetuning technique (Aim 1b). Aim 2: utilize zero-shot inference (Aim 2a) and soft-prompt tuning (Aim 2b) on these fine-tuned models to predict gastric cancer data elements. This proposal is innovatite by using LLMs to identify key cancer data elements in real-world settings, and has broad impacts by accelerating research, streamlining cancer registry operations, and fostering the development of effective cancer prevention and treatment therapies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Project 3
Project 3
Project 3
New Therapeutic Vulnerabilities for Aggressive B-Cell Lymphoma
海外基金