Fextor: A Feature Extraction Framework for Natural Language Processing: A Case Study in Word Sense Disambiguation, Relation Recognition and Anaphora Resolution

Fextor: A Feature Extraction Framework for Natural Language Processing: A Case Study in Word Sense Disambiguation, Relation Recognition and Anaphora Resolution
复制标题

DOI:
10.1007/978-3-642-34399-5_3
复制
发表时间:
2013-01-01
期刊:
COMPUTATIONAL LINGUISTICS: APPLICATIONS
影响因子:
--
通讯作者:
Wardynski, Adam
Wardynski, Adam
中科院分区:
其他
文献类型:
--
作者:
Broda, Bartosz;Kedzia, Pawel;Wardynski, Adam

文献摘要

被引文献

相似文献

从文本语料库中提取特征是自然语言处理(NLP),特别是机器学习(ML)技术的重要步骤。各种NLP任务有许多共同的步骤,例如阅读语料库并从中获取文本窗口的低级行为。一些高级处理步骤也可能是共享的,例如测试单词之间的形态句法约束。一个集成的特征提取框架消除了浪费的冗余,并有助于快速原型。在本文中,我们提出了一个灵活的特征提取框架Fextor。我们描述的特征提取过程中的假设,并提供软件架构的一般概述。这是伴随着在巨大不同的NLP任务的应用程序的例子。即,我们展示了Fextor的应用:词义消歧,识别组块间的句法关系,命名实体之间的语义关系,以及回指解析。
Feature extraction from text corpora is an important step in Natural Language Processing (NLP), especially for Machine Learning (ML) techniques. Various NLP tasks have many common steps, e.g. low level act of reading a corpus and obtaining text windows from it. Some high-level processing steps might also be shared, e.g. testing for morpho-syntactic constraints between words. An integrated feature extraction framework removes wasteful redundancy and helps in rapid prototyping.In this paper we present a flexible feature extraction framework called Fextor. We describe assumptions about the feature extraction process and provide general overview of software architecture. This is accompanied by examples of applications in hugely different NLP tasks. Namely, we show the application of Fextor in: word sense disambiguation, recognition of inter-chunk syntactic relations, semantic relations between named entities, as well as anaphora resolution.