A Systematic Exploration of the Feature Space for Relation Extraction

A Systematic Exploration of the Feature Space for Relation Extraction
复制标题

DOI:
--
复制
发表时间:
2007-12
期刊:
--
影响因子:
--
通讯作者:
Jing Jiang;ChengXiang Zhai
Jing Jiang;ChengXiang Zhai
中科院分区:
其他
文献类型:
--
作者:
Jing Jiang;ChengXiang Zhai

文献摘要

被引文献

相似文献

关系抽取是从文本中发现实体之间的语义关系的任务。最先进的关系提取方法大多基于统计学习,因此都必须处理特征选择,这会显着影响分类性能。在本文中,我们系统地探索了一个大的空间的特征关系提取和评估不同的特征子空间的有效性。我们提出了一个一般的定义的特征空间的基础上的关系实例的图形表示,并探讨了三种不同的表示关系实例和功能的不同复杂性在这个框架内。我们的实验表明,仅使用基本的单元特征通常足以实现最先进的性能,而过度包含复杂的特征可能会损害性能。不同复杂程度和不同句子表示的特征的组合,加上面向任务的特征修剪,给出了最佳性能。
Relation extraction is the task of finding semantic relations between entities from text. The state-of-the-art methods for relation extraction are mostly based on statistical learning, and thus all have to deal with feature selection, which can significantly affect the classification performance. In this paper, we systematically explore a large space of features for relation extraction and evaluate the effectiveness of different feature subspaces. We present a general definition of feature spaces based on a graphic representation of relation instances, and explore three different representations of relation instances and features of different complexities within this framework. Our experiments show that using only basic unit features is generally sufficient to achieve state-of-the-art performance, while overinclusion of complex features may hurt the performance. A combination of features of different levels of complexity and from different sentence representations, coupled with task-oriented feature pruning, gives the best performance.