A Neural Candidate-Selector Architecture for Automatic Structured Clinical Text Annotation.

A Neural Candidate-Selector Architecture for Automatic Structured Clinical Text Annotation.
复制标题

DOI:
10.1145/3132847.3132989
复制
发表时间:
2017-11
期刊:
Proceedings of the ... ACM International Conference on Information & Knowledge Management. ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Wallace BC
Wallace BC
中科院分区:
其他
文献类型:
--
作者:
Singh G;Marshall IJ;Thomas J;Shawe-Taylor J;Wallace BC

文献摘要

被引文献

相似文献

我们认为,自动注释自由文本描述的概念,从一个控制的,结构化的医学词汇的临床试验的任务。具体来说,我们的目标是建立一个模型来推断不同的(本体论)概念集,这些概念描述了基础试验的互补临床突出方面:招募的人群,管理的干预措施和测量的结果,即,皮科元素这一重要的实际问题带来了一些关键挑战。一个问题是输出空间很大,因为词汇表包含许多独特的概念。使这个问题更加复杂的是,在这个领域中的带注释的数据收集起来是昂贵的,因此是稀疏的。此外,输出(每个皮科元素的概念集)是相关的:特定人群(例如,糖尿病患者)将使某些干预概念成为可能(胰岛素治疗),同时有效地排除其他干预概念(放射治疗)。应该利用这种相关性。我们提出了一种新型神经模型来解决这些挑战。我们引入了一个可扩展的架构,在该模型中,考虑皮科元素的候选概念集,并评估其可扩展性的输入文本进行注释的条件。这依赖于“候选集”生成器,其可以被学习或依赖于算法。条件判别神经模型,然后联合选择候选概念,给定的输入文本。我们将我们的方法的预测性能与强基线进行比较,并表明它优于它们。最后,我们进行了定性评估所生成的注释,要求领域专家评估其质量。
We consider the task of automatically annotating free texts describing clinical trials with concepts from a controlled, structured medical vocabulary. Specifically we aim to build a model to infer distinct sets of (ontological) concepts describing complementary clinically salient aspects of the underlying trials: the populations enrolled, the interventions administered and the outcomes measured, i.e., the PICO elements. This important practical problem poses a few key challenges. One issue is that the output space is vast, because the vocabulary comprises many unique concepts. Compounding this problem, annotated data in this domain is expensive to collect and hence sparse. Furthermore, the outputs (sets of concepts for each PICO element) are correlated: specific populations (e.g., diabetics) will render certain intervention concepts likely (insulin therapy) while effectively precluding others (radiation therapy). Such correlations should be exploited. We propose a novel neural model that addresses these challenges. We introduce a Candidate-Selector architecture in which the model considers setes of candidate concepts for PICO elements, and assesses their plausibility conditioned on the input text to be annotated. This relies on a ‘candidate set’ generator, which may be learned or relies on heuristics. A conditional discriminative neural model then jointly selects candidate concepts, given the input text. We compare the predictive performance of our approach to strong baselines, and show that it outperforms them. Finally, we perform a qualitative evaluation of the generated annotations by asking domain experts to assess their quality.