Supervised line attention for tumor attribute classification from pathology reports: Higher performance with less data

Supervised line attention for tumor attribute classification from pathology reports: Higher performance with less data
复制标题

DOI:
10.1016/j.jbi.2021.103872
复制
发表时间:
2021-09-14
影响因子:
4.5
通讯作者:
Yu, Bin
Yu, Bin
中科院分区:
医学3区
文献类型:
--
作者:
Altieri, Nicholas;Park, Briton;Yu, Bin

文献摘要

被引文献

相似文献

目的:我们的目标是建立一个准确的基于机器学习的系统,用于在存在少量注释数据的情况下对癌症病理报告中的肿瘤属性进行分类,这是由于病理报告注释的昂贵和耗时。包括相关信息的位置沿着最终标签的丰富的标签方案与利用这些丰富的注释的用于对报告进行分类的对应的分层方法一起沿着使用。材料与方法:我们的数据包括2002年至2019年加州大学旧金山分校弗朗西斯科的250份结肠癌和250份肾癌病理报告。对于每份报告,我们对属性进行了分类,例如进行的手术、肿瘤分级和肿瘤部位。对于每个属性和文档,由肿瘤学家训练的注释器标记该属性的值以及文档中指示该值的特定行。我们开发了一个模型,该模型使用这些丰富的注释,首先预测文档的相关行,然后预测给定预测行的最终值。我们将我们的模型与多种最先进的方法进行比较,用于从病理报告中对肿瘤属性进行分类。结果如下:我们的研究结果表明,在结肠癌和肾癌以及不同的训练集大小中,我们的分层方法始终优于最先进的方法。此外,与这些方法相当的性能可以用大约一半的标记数据量来实现。结论:文档注释,丰富的位置信息,大大提高了机器学习方法的样本效率分类属性的病理报告。
Objective: We aim to build an accurate machine learning-based system for classifying tumor attributes from cancer pathology reports in the presence of a small amount of annotated data, motivated by the expensive and time-consuming nature of pathology report annotation. An enriched labeling scheme that includes the location of relevant information along with the final label is used along with a corresponding hierarchical method for classifying reports that leverages these enriched annotations. Materials and methods: Our data consists of 250 colon cancer and 250 kidney cancer pathology reports from 2002 to 2019 at the University of California, San Francisco. For each report, we classify attributes such as procedure performed, tumor grade, and tumor site. For each attribute and document, an annotator trained by an oncologist labeled both the value of that attribute as well as the specific lines in the document that indicated the value. We develop a model that uses these enriched annotations that first predicts the relevant lines of the document, then predicts the final value given the predicted lines. We compare our model to multiple state-of-the-art methods for classifying tumor attributes from pathology reports. Results: Our results show that across colon and kidney cancers and varying training set sizes, our hierarchical method consistently outperforms state-of-the-art methods. Furthermore, performance comparable to these methods can be achieved with approximately half the amount of labeled data. Conclusion: Document annotations that are enriched with location information are shown to greatly increase the sample efficiency of machine learning methods for classifying attributes of pathology reports.