Automatic Term Extraction Based on Perplexity of Compound Words

Automatic Term Extraction Based on Perplexity of Compound Words
复制标题

DOI:
10.1007/11562214_24
复制
发表时间:
2005-10
期刊:
--
影响因子:
--
通讯作者:
Minoru Yoshida;Hiroshi Nakagawa
Minoru Yoshida;Hiroshi Nakagawa
中科院分区:
其他
文献类型:
--
作者:
Minoru Yoshida;Hiroshi Nakagawa

文献摘要

相似文献

在大型语料库上,许多术语提取方法都是根据其准确性进行讨论的。然而,当我们试图将从频率派生的各种方法应用到一个较小的语料库时,由于缺乏关于频率的统计信息,我们可能无法达到足够的准确率。本文报告了一种新的术语提取方法,该方法针对非常小的语料库进行了调整。它侧重于复合词的结构,并计算了复合词单元左右两侧的困惑程度。实验结果表明,该方法的准确率并不是很高。然而,与其他方法相比,结合困惑和频率信息的方法的实验获得了最高的平均精度。
Many methods of term extraction have been discussed in terms of their accuracy on huge corpora. However, when we try to apply various methods that derive from frequency to a small corpus, we may not be able to achieve sufficient accuracy because of the shortage of statistical information on frequency. This paper reports a new way of extracting terms that is tuned for a very small corpus. It focuses on the structure of compound terms and calculates perplexity on the term unit’s left-side and right-side. The results of our experiments revealed that the accuracy with the proposed method was not that advantageous. However, experimentation with the method combining perplexity and frequency information obtained the highest average-precision in comparison with other methods.