ToKSA - Tokenized Key Sentence Annotation - a Novel Method for Rapid Approximation of Ground Truth for Natural Language Processing

ToKSA - Tokenized Key Sentence Annotation - a Novel Method for Rapid Approximation of Ground Truth for Natural Language Processing
复制标题

DOI:
10.1101/2021.10.06.21264629
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
C. Fairfield;W. Cambridge;L. Cullen;T. Drake;S. Knight;N. Masson;N. Mills;R. Pius;C. A. Shaw;H. Wu;S. Wigmore;A. Spiliopoulou;E. M. Harrison
C. Fairfield;W. Cambridge;L. Cullen;T. Drake;S. Knight;N. Masson;N. Mills;R. Pius;C. A. Shaw;H. Wu;S. Wigmore;A. Spiliopoulou;E. M. Harrison
中科院分区:
其他
文献类型:
--
作者:
C. Fairfield;W. Cambridge;L. Cullen;T. Drake;S. Knight;N. Masson;N. Mills;R. Pius;C. A. Shaw;H. Wu;S. Wigmore;A. Spiliopoulou;E. M. Harrison

文献摘要

相似文献

目的从自由文本中识别表型和病理是临床工作和研究的重要任务。自然语言处理(NLP)是大规模处理自由文本的关键工具。开发和验证NLP模型需要标记数据。标签是通过耗时和重复的手动注释生成的,并且对于敏感的临床数据很难获得。本文的目的是描述一种新的方法来注释放射学报告。材料和方法我们实现了标记化的关键药物特异性注释(ToKSA)来注释临床数据。我们使用180,050份腹部超声报告证明了ToKSA,这些报告具有针对症状状态、胆结石状态和胆囊切除术状态生成的标签。首先,将单个句子分组到词频矩阵中。然后使用关键(即最频繁出现的)句子的注释来同时为多个报告生成标签。我们将ToKSA衍生的标签与通过注释完整报告生成的标签进行了比较。我们使用ToKSA衍生的标签来训练使用卷积神经网络的文档分类器。我们将分类器的性能与基于完整报告的标签训练的单独分类器进行了比较。结果仅通过注释2,000个频繁句子,我们就能够为70,000份报告(准确率98.4%)、85,177份报告(准确率99.2%)和85,177份报告(准确率100%)生成症状状态标签。在ToKSA标签上训练的文档分类器的准确度与在完整报告标签上训练的文档分类器相似(准确度高出0.1-1.1%)。结论ToKSA是一种准确、高效的标注自由文本临床数据的方法。
Objective Identifying phenotypes and pathology from free text is an essential task for clinical work and research. Natural language processing (NLP) is a key tool for processing free text at scale. Developing and validating NLP models requires labelled data. Labels are generated through time-consuming and repetitive manual annotation and are hard to obtain for sensitive clinical data. The objective of this paper is to describe a novel approach for annotating radiology reports. Materials and Methods We implemented tokenized key sentence-specific annotation (ToKSA) for annotating clinical data. We demonstrate ToKSA using 180,050 abdominal ultrasound reports with labels generated for symptom status, gallstone status and cholecystectomy status. Firstly, individual sentences are grouped together into a term-frequency matrix. Annotation of key (i.e. the most frequently occurring) sentences is then used to generate labels for multiple reports simultaneously. We compared ToKSA-derived labels to those generated by annotating full reports. We used ToKSA-derived labels to train a document classifier using convolutional neural networks. We compared performance of the classifier to a separate classifier trained on labels based on the full reports. Results By annotating only 2,000 frequent sentences, we were able to generate labels for symptom status for 70,000 reports (accuracy 98.4%), gallstone status for 85,177 reports (accuracy 99.2%) and cholecystectomy status for 85,177 reports (accuracy 100%). The accuracy of the document classifier trained on ToKSA labels was similar (0.1-1.1% more accurate) to the document classifier trained on full report labels. Conclusion ToKSA offers an accurate and efficient method for annotating free text clinical data.