Automatic extraction of property norm-like data from large text corpora

Automatic extraction of property norm-like data from large text corpora
复制标题

从大型文本语料库中自动提取类似属性规范的数据

DOI:
--
复制
发表时间:
2014
期刊:
Cognitive Sciences
影响因子:
--
通讯作者:
Colin Kelly
Colin Kelly
中科院分区:
--
文献类型:
--
作者:
Colin Kelly

文献摘要

参考文献

被引文献

相似文献

从文本中派生基于属性的概念表示的传统方法要么专注于仅提取可能关系类型的子集,例如下名/上名(例如,汽车是一辆汽车)或名/转喻(例如,汽车有轮子),要么未指定关系(例如,汽车-汽油)。我们提出了一个系统,用于从大型文本语料库中自动、大规模地获取不受约束的类人属性规范,并讨论了该系统的理论含义。我们使用句法、语义和百科全书式的信息来指导我们的提取,产生概念-关系-特征三元组(例如,汽车快、汽车需要汽油、汽车造成污染),它们近似于基于属性的概念表示。我们的新方法使用语法和语法动机规则从已解析的语料库(维基百科和英国国家语料库)中提取候选三元组,然后使用它们的频率和四个统计指标的线性组合重新加权三元组。我们通过三种方式评估我们的系统输出:与人类生成的属性规范数据衍生的规范进行词汇比较,由四名人类评委直接评估,以及与WordNet相似度数据和人类判断的概念相似度评级进行语义距离比较。我们的系统提供了一种可行的、高性能的三重提取方法:我们的词法比较显示了与当前最先进的性能相当的性能,而随后的评估显示了我们生成的属性的类似人类的特征。
Traditional methods for deriving property-based representations of concepts from text have focused on either extracting only a subset of possible relation types, such as hyponymy/hypernymy (e.g., car is-a vehicle) or meronymy/metonymy (e.g., car has wheels), or unspecified relations (e.g., car--petrol). We propose a system for the challenging task of automatic, large-scale acquisition of unconstrained, human-like property norms from large text corpora, and discuss the theoretical implications of such a system. We employ syntactic, semantic, and encyclopedic information to guide our extraction, yielding concept-relation-feature triples (e.g., car be fast, car require petrol, car cause pollution), which approximate property-based conceptual representations. Our novel method extracts candidate triples from parsed corpora (Wikipedia and the British National Corpus) using syntactically and grammatically motivated rules, then reweights triples with a linear combination of their frequency and four statistical metrics. We assess our system output in three ways: lexical comparison with norms derived from human-generated property norm data, direct evaluation by four human judges, and a semantic distance comparison with both WordNet similarity data and human-judged concept similarity ratings. Our system offers a viable and performant method of plausible triple extraction: Our lexical comparison shows comparable performance to the current state-of-the-art, while subsequent evaluations exhibit the human-like character of our generated properties.
DOI: --
发表时间: 2008-07
期刊: --
影响因子: --
作者:
George S. Cree;K. McRae;Larry Barsalou;Mike Dixon;Jeff Elman;Albert Katz
通讯作者: George S. Cree;K. McRae;Larry Barsalou;Mike Dixon;Jeff Elman;Albert Katz
DOI: --
发表时间: 2000
期刊: --
影响因子: --
作者:
Dan Jurafsky;James H. Martin
通讯作者: Dan Jurafsky;James H. Martin
DOI: --
发表时间: 2010-06
期刊: --
影响因子: --
作者:
Barry Devereux;Colin Kelly;A. Korhonen
通讯作者: Barry Devereux;Colin Kelly;A. Korhonen
DOI: 10.1037/0096-3445.120.4.339
发表时间: 1991-12-01
影响因子: 4.1
作者:
FARAH, MJ;MCCLELLAND, JL
通讯作者: MCCLELLAND, JL
DOI: 10.1037//0096-3445.126.2.99
发表时间: 1997-06
期刊: Journal of experimental psychology. General
影响因子: --
作者:
K. McRae;V. D. de Sa;Mark S. Seidenberg
通讯作者: K. McRae;V. D. de Sa;Mark S. Seidenberg