Clustering Words with the MDL Principle
Clustering Words with the MDL Principle
复制标题
使用 MDL 原则对单词进行聚类
DOI:
10.3115/992628.992633
复制
发表时间:
1996
期刊:
影响因子:
--
通讯作者:
N. Abe
中科院分区:
文献类型:
--
作者:
Hang Li;N. Abe
We address the problem of automatically constructing a thesaurus by clustering words based on corpus data. We view this problem as that of estimating a joint distribution over the Cartesian product of a partition of a set of nouns and a partition of a set of verbs, and propose a learning algorithm based on the Minimum Description Length (MDL) Principle for such estimation. We empirically compared the performance of our method based on the MDL Principle against the Maximum Likelihood Estimator in word clustering, and found that the former outperforms the latter. We also evaluated the method by conducting pp-attachment disambiguation experiments using an automatically constructed thesaurus. Our experimental results indicate that such a thesaurus can be used to improve accuracy in disambiguation.