课题基金 / 基金详情

Applying Large Language Models to Accelerate Abstraction of Cancer Pathology Reports for Cancer Registry (LLMs for Unstructured Data Extraction)

Applying Large Language Models to Accelerate Abstraction of Cancer Pathology Reports for Cancer Registry (LLMs for Unstructured Data Extraction)
应用大型语言模型加速癌症登记的癌症病理报告的抽象(非结构化数据提取法学硕士)
批准号:
10890243
负责人:
John L. Cleveland
金额:
$30.0万
依托单位国家:
美国
项目类别:
财政年份:
1998
资助国家:
美国
项目状态:
未结题
起止时间:
1998-02-18 至 2027-01-31

项目摘要

项目成果

John L. Cleveland的其他基金

相似基金

相关文献

中文摘要
翻译
病理学报告包含组织样本和病变的关键信息,在以下方面发挥着重要作用: 确定癌症治疗选择、预后、风险分层和临床试验筛选。然而,手动 从这些非结构化或半结构化的报告中提取肿瘤特征是一个复杂、费力的过程。 过程通过深度学习方法进行自然语言处理(NLP)的最新进展显示, 很有潜力虽然双向编码器表示从变压器(BERT)已经实现了 尽管在各种NLP任务中取得了显著的成绩,但由于允许的有限性,其在病理学中的应用受到限制 输入长度。我们最近的研究通过迁移学习解决了这一问题,迁移学习是一种基于BERT的模型, 复杂的知识来源,包括维基百科,PubMed,MIMIC-III和Moffitt机构病理学 报道进一步微调该语言模型,以识别部位、组织学和相关ICD-O-3代码 从病理报告中。尽管初步结果令人鼓舞,但我们的试点工作主要集中在采掘业。 单一原发性实体瘤诊断的问题回答,忽略了丰富的术语和变化, 病理学语言 我们的长期目标是采用大型语言模型(LLM)从所有类型的临床信息中提取信息。 注意到,协助机构认证的肿瘤登记在数据提取癌症登记。在这项工作中, 我们特别关注病理学报告,并建议对349,544名机构病理学的法学硕士进行培训 报告,以确定五个关键的癌症数据元素:原发部位,组织学,阶段,等级和偏侧性。研究 将重点关注常见(乳腺癌)和罕见(胃癌)癌症。 我们将利用现有的LLM在大型公共语料库上预先培训, 报告,并最终微调它们以预测特定的癌症数据元素。 我们追求两个具体目标。目标1:通过抽象问题询问预测乳腺癌数据元素 使用现有的解百纳架构(Aim 1a),并通过基于葡萄酒的微调技术(Aim 1b)。目标二: 在这些微调过的模型上利用零触发推理(Aim 2a)和软提示调整(Aim 2b)来预测 胃癌数据元素。该提案是创新的,通过使用LLM来识别关键的癌症数据元素, 现实世界的环境,并通过加速研究,简化癌症登记操作, 以及促进有效的癌症预防和治疗疗法的发展。
英文摘要
Pathology reports, containing critical information on tissue samples and lesions, play a significant role in determining cancer treatment selection, prognosis, risk stratification, and clinical trial screening. Yet, manually extracting tumor characteristics from these unstructured or semi-structured reports is a complex, laborious process. Recent advances in Natural Language Processing (NLP) via deep learning methodologies show promising potential. Though Bidirectional Encoder Representations from Transformers (BERT) has achieved notable results in various NLP tasks, its application in pathology is constrained due to the limited allowable input length. Our recent study addressed this by transfer learning a BERT-based model on increasingly complex knowledge sources including Wikipedia, PubMed, MIMIC-III, and Moffitt institutional pathology reports. This language model was further fine-tuned to identify site, histology, and associated ICD-O-3 codes from pathology reports. Despite the promising preliminary results, our pilot work focuses on extractive question-anwsering of single primary solid tumor diagnosis, overlooking rich terminology and variation of the pathology language. Our long-term goal is to employ Large Language Models (LLMs) to extract information from all types of clinical notes, assisting institutional certified tumor registrars in data abstraction for the Cancer Registry. In this work, we specifically focus on pathology reports, and proposes to train LLMs on 349,544 institutional pathology reports to identify five key cancer data elements: primary site, histology, stage, grade, and laterality. The study will focus on common (breast) and rare (gastric) cancers. We will leverage existing LLMs pretrained on large public corpora, retrain them on institutional pathology reports, and finally fine-tune them to predict specific cancer data elements. We pursue two specific aims. Aim 1: predict breast cancer data elements by abstractive question-aswering using the existing cabernet architecture (Aim 1a), and by a prompt-based finetuning technique (Aim 1b). Aim 2: utilize zero-shot inference (Aim 2a) and soft-prompt tuning (Aim 2b) on these fine-tuned models to predict gastric cancer data elements. This proposal is innovatite by using LLMs to identify key cancer data elements in real-world settings, and has broad impacts by accelerating research, streamlining cancer registry operations, and fostering the development of effective cancer prevention and treatment therapies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Project 3
Project 3
Project 3
New Therapeutic Vulnerabilities for Aggressive B-Cell Lymphoma
海外基金