Clustering Words with the MDL Principle

Clustering Words with the MDL Principle
复制标题

使用 MDL 原则对单词进行聚类

DOI:
10.3115/992628.992633
复制
发表时间:
1996
期刊:
ArXiv
影响因子:
--
通讯作者:
N. Abe
N. Abe
中科院分区:
--
文献类型:
--
作者:
Hang Li;N. Abe

文献摘要

被引文献

相似文献

我们解决的问题,自动构建一个基于语料库数据的聚类词的词库。我们认为这个问题,估计一组名词的分区和一组动词的分区的笛卡尔积的联合分布,并提出了一个学习算法的基础上的最小描述长度(MDL)的原则,这样的估计。我们实证比较了我们的方法的性能基于MDL原则的最大似然估计在词聚类,发现前者优于后者。我们还评估了该方法进行PP附件消歧实验,使用自动构建的词库。我们的实验结果表明,这样的词库可以用来提高准确率在消歧。
We address the problem of automatically constructing a thesaurus by clustering words based on corpus data. We view this problem as that of estimating a joint distribution over the Cartesian product of a partition of a set of nouns and a partition of a set of verbs, and propose a learning algorithm based on the Minimum Description Length (MDL) Principle for such estimation. We empirically compared the performance of our method based on the MDL Principle against the Maximum Likelihood Estimator in word clustering, and found that the former outperforms the latter. We also evaluated the method by conducting pp-attachment disambiguation experiments using an automatically constructed thesaurus. Our experimental results indicate that such a thesaurus can be used to improve accuracy in disambiguation.