Automatic Extraction of Fixed Multiword Expressions

Automatic Extraction of Fixed Multiword Expressions
复制标题

DOI:
10.1007/11562214_50
复制
发表时间:
2005-05
期刊:
--
影响因子:
--
通讯作者:
Campbell Hore;Masayuki Asahara;Yuji Matsumoto
Campbell Hore;Masayuki Asahara;Yuji Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Campbell Hore;Masayuki Asahara;Yuji Matsumoto

文献摘要

被引文献

相似文献

固定的多词表达式是一串词,它们一起表现得像一个词。本研究建立了一种自动提取此类表达式的方法。我们的方法包括三个阶段。在第一种方法中,使用统计测量来提取候选二元组。在第二种情况下,我们使用这个列表来选择出现的候选表达式在语料库中,连同他们周围的上下文。这些示例用作监督机器学习的训练数据,从而产生可以识别目标多词表达的分类器。最后一个阶段是估计的语音部分的每一个提取的表达的基础上,其上下文的发生。评估表明,搭配措施单独是不能有效地识别目标表达式。然而,当在一百万个例子上训练时,分类器识别目标多词表达的精度大于90%。词性估计的准确率和召回率均在95%以上。
Fixed multiword expressions are strings of words which together behave like a single word. This research establishes a method for the automatic extraction of such expressions. Our method involves three stages. In the first, a statistical measure is used to extract candidate bigrams. In the second, we use this list to select occurrences of candidate expressions in a corpus, together with their surrounding contexts. These examples are used as training data for supervised machine learning, resulting in a classifier which can identify target multiword expressions. The final stage is the estimation of the part of speech of each extracted expression based on its context of occurence. Evaluation demonstrated that collocation measures alone are not effective in identifying target expressions. However, when trained on one million examples, the classifier identified target multiword expressions with precision greater than 90%. Part of speech estimation had precision and recall of over 95%.