Benchmark Pathology Report Text Corpus with Cancer Type Classification.

Benchmark Pathology Report Text Corpus with Cancer Type Classification.
复制标题

具有癌症类型分类的基准病理学报告文本语料库。

DOI:
10.1101/2023.08.03.23293618
复制
发表时间:
2023
期刊:
medRxiv : the preprint server for health sciences
影响因子:
--
通讯作者:
Tatonetti,Nicholas
Tatonetti,Nicholas
中科院分区:
--
文献类型:
--
作者:
Kefeli,Jenna;Tatonetti,Nicholas

文献摘要

被引文献

相似文献

在癌症研究中,病理报告文本在很大程度上是一个未开发的数据源。病理报告是常规生成的,比结构化数据更细致,并包含病理学家的见解。然而,没有公开可用的数据集来对基于报告的模型进行基准测试。最近的两项进展表明,迫切需要一个基准数据集。首先,改进的光学字符识别(OCR)技术将使以自动化的方式访问旧的病理报告成为可能,增加了可用于分析的数据。其次,最近使用人工智能的自然语言处理(NLP)技术的改进可以从文本中更准确地预测临床目标。我们将最先进的OCR和定制的后处理应用于癌症基因组图谱中公开可用的报告pdf,生成9,523份报告的机器可读语料库。我们在32个组织中进行了原则性的癌症类型分类,平均AU-ROC达到0.992。该数据集将对各个专业的研究人员有用,包括研究临床医生、临床试验研究者和临床NLP研究人员。
In cancer research, pathology report text is a largely un-tapped data source. Pathology reports are routinely generated, more nuanced than structured data, and contain added insight from pathologists. However, there are no publicly-available datasets for benchmarking report-based models. Two recent advances suggest the urgent need for a benchmark dataset. First, improved optical character recognition (OCR) techniques will make it possible to access older pathology reports in an automated way, increasing data available for analysis. Second, recent improvements in natural language processing (NLP) techniques using AI allow more accurate prediction of clinical targets from text. We apply state-of-the-art OCR and customized post-processing to publicly available report PDFs from The Cancer Genome Atlas, generating a machine-readable corpus of 9,523 reports. We perform a proof-of-principle cancer-type classification across 32 tissues, achieving 0.992 average AU-ROC. This dataset will be useful to researchers across specialties, including research clinicians, clinical trial investigators, and clinical NLP researchers.