A natural language processing pipeline for pairing measurements uniquely across free-text CT reports

A natural language processing pipeline for pairing measurements uniquely across free-text CT reports
复制标题

DOI:
10.1016/j.jbi.2014.08.015
复制
发表时间:
2015-02-01
影响因子:
4.5
通讯作者:
Trost, William
Trost, William
中科院分区:
医学3区
文献类型:
--
作者:
Sevenster, Merlijn;Bozeman, Jeffrey;Trost, William

文献摘要

被引文献

相似文献

目的:为了使肿瘤学中的治疗反应评估标准化和客观化,已经提出了由放射学测量驱动的指南,这些测量通常在无视自动处理的自由文本报告中传达。我们研究通过inter-annotator协议和自然语言处理(NLP)算法开发的配对测量的任务,量化跨连续的放射学报告的相同发现,使每个测量是配对的最多一个其他(“部分唯一性”)。方法和材料:地面真相创建的基础上,283腹部和311胸部CT报告的50例患者。预处理引擎将报告分段并提取测量结果。十三个功能的基础上开发的测量之间的体积相似性,语义相似性各自的叙述上下文和结构属性的报告位置。随机森林分类器(RF)集成了所有特征。结果:在端到端评估中,RF的精确度为0.841,召回率为0.807,F-measure为0.824,AUC为0.971; MBM的精确度为0.899,召回率为0.776,F-measure为0.833,AUC为0.935,高于机会水平(P < 0.001)。RF(RF + MBM)在52.7%(57.4%)的报告pairs.Discussion上具有无错误性能:三个领域专家与地面实况(kappa > 0.960)的注释者间一致性表明任务定义良好。域属性和区间差异进行了讨论,以解释腹部的上级性能。强制执行部分的唯一性混合,但对performance.Conclusion轻微的影响:一个组合的机器学习过滤方法,提出了配对测量,它可以支持前瞻性(支持治疗反应评估)和回顾性的目的(数据挖掘)。(C)2014爱思唯尔公司All rights reserved.
Objective: To standardize and objectivize treatment response assessment in oncology, guidelines have been proposed that are driven by radiological measurements, which are typically communicated in free-text reports defying automated processing. We study through inter-annotator agreement and natural language processing (NLP) algorithm development the task of pairing measurements that quantify the same finding across consecutive radiology reports, such that each measurement is paired with at most one other ("partial uniqueness").Methods and materials: Ground truth is created based on 283 abdomen and 311 chest CT reports of 50 patients each. A pre-processing engine segments reports and extracts measurements. Thirteen features are developed based on volumetric similarity between measurements, semantic similarity between their respective narrative contexts and structural properties of their report positions. A Random Forest classifier (RF) integrates all features. A "mutual best match" (MBM) post-processor ensures partial uniqueness.Results: In an end-to-end evaluation, RF has precision 0.841, recall 0.807, F-measure 0.824 and AUC 0.971; with MBM, which performs above chance level (P < 0.001), it has precision 0.899, recall 0.776, F-measure 0.833 and AUC 0.935. RF (RF + MBM) has error-free performance on 52.7% (57.4%) of report pairs.Discussion: Inter-annotator agreement of three domain specialists with the ground truth (kappa > 0.960) indicates that the task is well defined. Domain properties and inter-section differences are discussed to explain superior performance in abdomen. Enforcing partial uniqueness has mixed but minor effects on performance.Conclusion: A combined machine learning-filtering approach is proposed for pairing measurements, which can support prospective (supporting treatment response assessment) and retrospective purposes (data mining). (C) 2014 Elsevier Inc. All rights reserved.