Automatically extracting cancer disease characteristics from pathology reports into a Disease Knowledge Representation Model

Automatically extracting cancer disease characteristics from pathology reports into a Disease Knowledge Representation Model
复制标题

DOI:
10.1016/j.jbi.2008.12.005
复制
发表时间:
2009-10-01
影响因子:
4.5
通讯作者:
de Groen, Piet C.
de Groen, Piet C.
中科院分区:
医学3区
文献类型:
--
作者:
Coden, Anni;Savova, Guergana;de Groen, Piet C.

文献摘要

被引文献

相似文献

我们引入了一个可扩展和可修改的知识表示模型来以可比较和一致的方式表示癌症疾病特征。我们描述了一个系统MedTAS/P,它从自由文本的病理报告中自动实例化知识表示模型。MedTAS/P基于开源框架,其组件使用自然语言处理原理、机器学习和规则来发现和填充模型的元素。为了验证模型并测量MedTAS/P的准确性,我们开发了手动注释的结肠癌病理报告的黄金标准语料库。对于知识表示模型中的实例化类,例如组织学或解剖位置,MedTAS/P的F1得分为0.97-1.0,对于需要提取关系的原发肿瘤或淋巴结,F1-得分为0.82-0.93。据报道,转移性肿瘤的F1得分为0.65,较低的得分主要是因为训练和测试集中的实例数量非常少。(C)2009 Elsevier Inc.保留所有权利。
We introduce an extensible and modifiable knowledge representation model to represent cancer disease characteristics in a comparable and consistent fashion. We describe a system, MedTAS/P which automatically instantiates the knowledge representation model from free-text pathology reports. MedTAS/P is based on an open-source framework and its components use natural language processing principles, machine learning and rules to discover and populate elements of the model. To validate the model and measure the accuracy of MedTAS/P, we developed a gold-standard corpus of manually annotated colon cancer pathology reports. MedTAS/P achieves F1-scores of 0.97-1.0 for instantiating classes in the knowledge representation model such as histologies or anatomical sites, and F1-scores of 0.82-0.93 for primary tumors or lymph nodes, which require the extractions of relations. An F1-score of 0.65 is reported for metastatic tumors, a lower score predominantly due to a very small number of instances in the training and test sets. (C) 2009 Elsevier Inc. All rights reserved.