A WordNet-based semantic approach to textual entailment and cross-lingual textual entailment

A WordNet-based semantic approach to textual entailment and cross-lingual textual entailment
复制标题

DOI:
10.1007/s13042-011-0026-z
复制
发表时间:
2011-07
影响因子:
5.6
通讯作者:
Julio J. Castillo
Julio J. Castillo
中科院分区:
计算机科学3区
文献类型:
--
作者:
Julio J. Castillo

文献摘要

被引文献

相似文献

本文阐述了如何在WordNet的基础上构建一个只使用语义相似性度量的文本蕴涵识别系统。我们展示了如何将广泛使用的基于WordNet的语义度量推广到构建句子级语义度量,以便同时用于单语和跨语言的文本蕴涵。我们用大量的RTE数据集进行了实验,并评估了一种扩展RTE单语言语料库的算法的贡献。用这种方法得到的结果在预测RTE测试集时产生了显著的统计差异。我们对这些度量进行了效率分析,得出了一些关于它们在识别文本蕴涵方面的实用价值的结论。我们还分析了跨语言的文本蕴涵任务,创建了一个英语-西班牙语双语语料库,并提出了一个为任意两种语言创建跨语言的文本蕴涵语料库的步骤。最后,我们证明了该方法足以构建一个单语和跨语言文本蕴涵的平均分数RTE系统,该系统使用来自WordNet的语义信息作为词汇语义知识的唯一来源。
In this paper we explain how to build a recognizing textual entailment (RTE) system which only uses semantic similarity measures based on WordNet. We show how the widely used WordNet-based semantic measures can be generalized to build sentence level semantic metrics in order to be used in both mono-lingual and cross-lingual textual entailment. We experiment with a wide variety of RTE datasets and evaluate the contribution of an algorithm which expands the RTE monolingual corpus. Results achieved with this method yielded significant statistical differences when predicting RTE test sets. We provide an efficiency analysis of these metrics drawing some conclusions about their practical utility in recognizing textual entailment. We also analyze the cross-lingual textual entailment task, we create a bilingual English–Spanish corpus, and propose a procedure to create a cross-lingual textual entailment corpus for any pair of languages. Finally, we show that the proposed method is enough to build an average score RTE system in both monolingual and cross-lingual textual entailment, that uses semantic information from WordNet as the only source of lexical-semantic knowledge.