From N-Grams to Collocations: An Evaluation of Xtract

From N-Grams to Collocations: An Evaluation of Xtract
复制标题

从 N-Gram 到搭配:Xtract 的评估

DOI:
10.3115/981344.981380
复制
发表时间:
1991
期刊:
--
影响因子:
--
通讯作者:
Frank Smadja
Frank Smadja
中科院分区:
--
文献类型:
--
作者:
Frank Smadja

文献摘要

被引文献

相似文献

在以前的论文中,我们提出了从大样本文本中检索搭配的方法。我们描述了一个工具,Xtract,实现这些方法,并能够检索范围广泛的搭配在两个阶段的过程。然而,这些方法以及其他相关方法具有一些局限性。主要是产生的搭配不包含任何功能信息,其中许多是无效的。在本文中,我们将介绍解决这些问题的方法。这些方法是在Xtract的第三阶段中实现的,该阶段检查在前两个阶段中检索到的搭配集,以过滤掉一些无效的搭配,并向保留的搭配添加有用的语法信息。通过结合解析和统计技术,第三阶段的加入使Xtract的整体精度水平从40%提高到80%,精度达到94%。在本文中,我们描述的方法和评价实验。
In previous papers we presented methods for retrieving collocations from large samples of texts. We described a tool, Xtract, that implements these methods and able to retrieve a wide range of collocations in a two stage process. These methods as well as other related methods however have some limitations. Mainly, the produced collocations do not include any kind of functional information and many of them are invalid. In this paper we introduce methods that address these issues. These methods are implemented in an added third stage to Xtract that examines the set of collocations retrieved during the previous two stages to both filter out a number of invalid collocations and add useful syntactic information to the retained ones. By combining parsing and statistical techniques the addition of this third stage has raised the overall precision level of Xtract from 40% to 80% with a precision of 94%. In the paper we describe the methods and the evaluation experiments.