Semantic Analysis and Automatic Corpus Construction for Entailment Recognition in Medical Texts

Semantic Analysis and Automatic Corpus Construction for Entailment Recognition in Medical Texts
复制标题

医学文本蕴涵识别的语义分析和自动语料库构建

DOI:
--
复制
发表时间:
2015
期刊:
Conference on Artificial Intelligence in Medicine in Europe
影响因子:
--
通讯作者:
Yassine Mrabet
Yassine Mrabet
中科院分区:
--
文献类型:
--
作者:
Asma Ben Abacha;Duy Dinh;Yassine Mrabet

文献摘要

被引文献

相似文献

文本蕴涵识别就是检测自然语言句子之间的推理关系。它有着广泛的应用,如机器翻译、问答或文本摘要。人们对RTE产生了浓厚的兴趣,但也面临着一些挑战。然而,目前的大多数方法都是专门针对开放领域的。技术和技术教育在专门领域面临的主要挑战是缺乏相关的培训语料库和资源。本文提出了一种用于医学领域RTE的语料库自动构建方法。我们还量化了使用(开放)域RDF数据集对基于RTE的监督学习的影响。我们通过比较基于记忆的高效学习算法在Pascal RTE语料库和我们自动构建的语料库上的结果来评估我们的语料库构建方法的相关性。结果表明,F测量的准确率提高了+6~+28%,提高了+8~+23%。我们还发现,来自大型开放领域数据集的语义标注将F1分数提高了6%,而较小的医疗RDF数据集实际上降低了整体性能。我们对这些发现进行了讨论,并对未来的研究提出了一些建议。
Textual Entailment Recognition (RTE) consists in detecting inference relationships between natural language sentences. It has a wide range of applications such as machine translation, question answering or text summarization. Significant interest has been brought to RTE with several challenges. However, most of current approaches are dedicated to open domains. The major challenge facing RTE in specialized domains is the lack of relevant training corpora and resources. In this paper we present an automatic corpus construction approach for RTE in the medical domain. We also quantify the impact of using (open-)domain RDF datasets on supervised learning based RTE. We evaluate the relevance of our corpus construction method by comparing the results obtained by an efficient memory based learning algorithm on PASCAL RTE corpora and on our automatically constructed corpus. The results show an accuracy increase of +6 to +28% and an improvement of +8 to +23% in terms of F-measure. We also found that semantic annotations from large open-domain datasets increased F1 score by 6%, while smaller medical RDF datasets actually decreased the overall performance. We discuss these findings and give some pointers to future investigations.