Automated ancillary cancer history classification for mesothelioma patients from free-text clinical reports.

Automated ancillary cancer history classification for mesothelioma patients from free-text clinical reports.
复制标题

DOI:
10.4103/2153-3539.71065
复制
发表时间:
2010-10-11
影响因子:
--
通讯作者:
Chapman BE
Chapman BE
中科院分区:
其他
文献类型:
--
作者:
Wilson RA;Chapman WW;Defries SJ;Becich MJ;Chapman BE

文献摘要

被引文献

相似文献

临床记录通常是非结构化的自由文本文档,这给信息提取带来了挑战和成本。医疗保健提供和研究组织,如国家间皮瘤虚拟银行,需要聚合结构化和非结构化数据类型。自然语言处理提供了从非结构化的自由文本文档中自动提取信息的技术。来自间皮瘤患者的508份病史和身体报告被分为开发(208)和测试集(300)。制定了一个参考标准,每一份报告都由专家根据患者的辅助癌症个人病史和任何癌症的家族史进行注释。开发HX应用程序是为了处理报告、提取相关特征、执行参考解析并根据癌症病史对其进行分类。对动态窗口和上下文两种信息提取方法进行了评价。对照参考标准测量使用这两种方法的HX的分类反应。科恩的平均加权卡帕作为评估该系统的人类基准。HX的总体准确率较高,每种方法的准确率为96.2%。使用动态窗口和上下文方法的F-测量分别为91.8%和91.6%,与人类基准的92.8%具有可比性。对于个人病史分类,动态窗口得分最高,为89.2%,对于家族史分类,背景得分最高,为97.6%,这两种方法与人类基准分别为88.3%和97.2%。我们评估了一个自动化应用程序在从临床报告中对间皮瘤患者的个人和家族癌症病史进行分类的性能。要做到这一点,HX应用程序必须处理报告,识别癌症概念,区分已知的间皮瘤和辅助癌症,识别否定,执行参考解析并确定体验者。结果表明,两种信息提取方法都依赖于领域词典和否定词的提取。我们展示了更通用的方法、上下文,以及我们的特定于任务的方法。虽然动态窗口可以被修改来检索其他概念,但上下文更健壮,在不确定的概念上表现得更好。HX可以极大地改进和加快从各种研究或医疗保健提供组织的自由文本临床记录中提取数据的过程。
Clinical records are often unstructured, free-text documents that create information extraction challenges and costs. Healthcare delivery and research organizations, such as the National Mesothelioma Virtual Bank, require the aggregation of both structured and unstructured data types. Natural language processing offers techniques for automatically extracting information from unstructured, free-text documents. Five hundred and eight history and physical reports from mesothelioma patients were split into development (208) and test sets (300). A reference standard was developed and each report was annotated by experts with regard to the patient’s personal history of ancillary cancer and family history of any cancer. The Hx application was developed to process reports, extract relevant features, perform reference resolution and classify them with regard to cancer history. Two methods, Dynamic-Window and ConText, for extracting information were evaluated. Hx’s classification responses using each of the two methods were measured against the reference standard. The average Cohen’s weighted kappa served as the human benchmark in evaluating the system. Hx had a high overall accuracy, with each method, scoring 96.2%. F-measures using the Dynamic-Window and ConText methods were 91.8% and 91.6%, which were comparable to the human benchmark of 92.8%. For the personal history classification, Dynamic-Window scored highest with 89.2% and for the family history classification, ConText scored highest with 97.6%, in which both methods were comparable to the human benchmark of 88.3% and 97.2%, respectively. We evaluated an automated application’s performance in classifying a mesothelioma patient’s personal and family history of cancer from clinical reports. To do so, the Hx application must process reports, identify cancer concepts, distinguish the known mesothelioma from ancillary cancers, recognize negation, perform reference resolution and determine the experiencer. Results indicated that both information extraction methods tested were dependant on the domain-specific lexicon and negation extraction. We showed that the more general method, ConText, performed as well as our task-specific method. Although Dynamic- Window could be modified to retrieve other concepts, ConText is more robust and performs better on inconclusive concepts. Hx could greatly improve and expedite the process of extracting data from free-text, clinical records for a variety of research or healthcare delivery organizations.