Automatically Detecting Failures in Natural Language Processing Tools for Online Community Text.

Automatically Detecting Failures in Natural Language Processing Tools for Online Community Text.
复制标题

DOI:
10.2196/jmir.4612
复制
发表时间:
2015-08-31
影响因子:
7.4
通讯作者:
Pratt W
Pratt W
中科院分区:
医学2区
文献类型:
--
作者:
Park A;Hartzler AL;Huh J;McDonald DW;Pratt W

文献摘要

被引文献

相似文献

患者生成的健康文本的普及率和价值正在增加,但处理此类文本仍然存在问题。尽管现有的生物医学自然语言处理(NLP)工具很有吸引力,但大多数工具都是为处理临床医生或研究人员生成的文本而开发的,例如临床笔记或期刊文章。除了为不同类型的文本构建之外,使用现有自然语言处理的其他挑战还包括不断变化的技术、源词汇和文本的特征。这些不断变化的挑战证明有必要应用低成本的系统评估。然而,NLP中最初被接受的评估方法-手动注释-需要付出巨大的努力和时间。这项研究的主要目标是探索一种替代方法-使用低成本的自动化方法来检测在使用现有生物医学NLP工具处理患者生成的文本时的故障(例如,错误的边界、遗漏的术语、错误映射的概念)。我们首先描述了NLP工具在处理在线社区文本时可能出现的常见故障。然后,我们使用最流行的生物医学NLP工具之一MetaMap演示了我们的自动化方法在检测这些常见故障方面的可行性。使用来自在线癌症社区的9657篇帖子,我们分两个步骤探索了我们的自动故障检测方法:(1)为了表征故障类型,我们首先手动审查MetaMap的常见故障,将不准确的映射归类为故障类型,然后使用开放编码通过迭代几轮手动审查来确定故障原因,以及(2)为了自动检测这些故障类型,我们随后探索了现有NLP技术和每个故障原因的基于词典的匹配的组合。最后,对自动检测到的故障进行了手动评估。从我们的手工回顾中,我们描述了三种类型的失误:(1)边界失误,(2)遗漏术语失误,和(3)词语歧义失误。在这三种失败类型中,我们发现了概念映射不准确的12个原因。我们使用自动化方法检测了385,572个MetaMap映射中的近一半存在问题。词义歧义失误的发生率最高,占82.22%。边界故障是第二常见的故障,占故障的15.90%,而遗漏项故障最不常见,占故障的1.88%。自动故障检测的准确率、召回率、准确率和F1得分分别为83.00%、92.57%、88.17%和87.52%。我们说明了处理患者生成的在线健康社区文本的挑战,并描述了NLP工具在该患者生成的健康文本上的故障特征,展示了我们自动检测这些故障的低成本方法的可行性。我们的方法显示了可扩展和有效的解决方案的潜力,以自动评估不断发展的NLP工具和源词汇表来处理患者生成的文本。
The prevalence and value of patient-generated health text are increasing, but processing such text remains problematic. Although existing biomedical natural language processing (NLP) tools are appealing, most were developed to process clinician- or researcher-generated text, such as clinical notes or journal articles. In addition to being constructed for different types of text, other challenges of using existing NLP include constantly changing technologies, source vocabularies, and characteristics of text. These continuously evolving challenges warrant the need for applying low-cost systematic assessment. However, the primarily accepted evaluation method in NLP, manual annotation, requires tremendous effort and time. The primary objective of this study is to explore an alternative approach—using low-cost, automated methods to detect failures (eg, incorrect boundaries, missed terms, mismapped concepts) when processing patient-generated text with existing biomedical NLP tools. We first characterize common failures that NLP tools can make in processing online community text. We then demonstrate the feasibility of our automated approach in detecting these common failures using one of the most popular biomedical NLP tools, MetaMap. Using 9657 posts from an online cancer community, we explored our automated failure detection approach in two steps: (1) to characterize the failure types, we first manually reviewed MetaMap’s commonly occurring failures, grouped the inaccurate mappings into failure types, and then identified causes of the failures through iterative rounds of manual review using open coding, and (2) to automatically detect these failure types, we then explored combinations of existing NLP techniques and dictionary-based matching for each failure cause. Finally, we manually evaluated the automatically detected failures. From our manual review, we characterized three types of failure: (1) boundary failures, (2) missed term failures, and (3) word ambiguity failures. Within these three failure types, we discovered 12 causes of inaccurate mappings of concepts. We used automated methods to detect almost half of 383,572 MetaMap’s mappings as problematic. Word sense ambiguity failure was the most widely occurring, comprising 82.22% of failures. Boundary failure was the second most frequent, amounting to 15.90% of failures, while missed term failures were the least common, making up 1.88% of failures. The automated failure detection achieved precision, recall, accuracy, and F1 score of 83.00%, 92.57%, 88.17%, and 87.52%, respectively. We illustrate the challenges of processing patient-generated online health community text and characterize failures of NLP tools on this patient-generated health text, demonstrating the feasibility of our low-cost approach to automatically detect those failures. Our approach shows the potential for scalable and effective solutions to automatically assess the constantly evolving NLP tools and source vocabularies to process patient-generated text.