Learning Syntactic Patterns for Automatic Hypernym Discovery

Learning Syntactic Patterns for Automatic Hypernym Discovery
复制标题

DOI:
--
复制
发表时间:
2004-12
影响因子:
--
通讯作者:
R. Snow;Dan Jurafsky;A. Ng
R. Snow;Dan Jurafsky;A. Ng
中科院分区:
计算机科学4区
文献类型:
--
作者:
R. Snow;Dan Jurafsky;A. Ng

文献摘要

被引文献

相似文献

语义分类法(如WordNet)为自然语言处理应用程序提供了丰富的知识来源,但构建、维护和扩展成本很高。出于自动构建和扩展此类分类的问题,本文提出了一种新的自动学习上位词(is-a)关系的算法。我们的方法概括了早期的工作,依赖于使用少量的手工制作的正则表达式模式,以确定上位词对。使用从解析树中提取的“依赖路径”特征,我们引入了这些模式的通用形式化和泛化。给定一个包含已知上位词对的文本训练集,我们的算法自动提取有用的依赖路径,并将其应用到新的语料库中以识别新的上位词对。在我们的评估任务(确定新闻文章中的两个名词是否参与上位词关系),我们自动提取的上位词数据库达到更高的精度和更高的召回率比WordNet。
Semantic taxonomies such as WordNet provide a rich source of knowledge for natural language processing applications, but are expensive to build, maintain, and extend. Motivated by the problem of automatically constructing and extending such taxonomies, in this paper we present a new algorithm for automatically learning hypernym (is-a) relations from text. Our method generalizes earlier work that had relied on using small numbers of hand-crafted regular expression patterns to identify hypernym pairs. Using "dependency path" features extracted from parse trees, we introduce a general-purpose formalization and generalization of these patterns. Given a training set of text containing known hypernym pairs, our algorithm automatically extracts useful dependency paths and applies them to new corpora to identify novel pairs. On our evaluation task (determining whether two nouns in a news article participate in a hypernym relationship), our automatically extracted database of hypernyms attains both higher precision and higher recall than WordNet.