Measuring historical word sense variation

Measuring historical word sense variation
复制标题

测量历史词义变化

DOI:
--
复制
发表时间:
2011
期刊:
ACM/IEEE Joint Conference on Digital Libraries
影响因子:
--
通讯作者:
G. Crane
G. Crane
中科院分区:
--
文献类型:
--
作者:
David Bamman;G. Crane

文献摘要

被引文献

相似文献

我们在这里描述了一种自动识别大型数字图书馆中过时的历史书籍集中的词义变化的方法。通过利用一小组已知的翻译书籍对来生成双语词义库存和 WSD 分类器的标记训练数据,我们能够自动对 3.89 亿个单词语料库中的拉丁词词义进行分类,并跟踪这些词义在 2000 年间的兴衰。我们在对对齐平行语料库中的 83,892 个单词和较小的手动注释的 525 个单词样本进行十倍测试中评估了七个不同分类器的性能,测量每个系统的整体准确性以及准确性(通过均方误差)与观察到的历史变化的相关程度。
We describe here a method for automatically identifying word sense variation in a dated collection of historical books in a large digital library. By leveraging a small set of known translation book pairs to induce a bilingual sense inventory and labeled training data for a WSD classifier, we are able to automatically classify the Latin word senses in a 389 million word corpus and track the rise and fall of those senses over a span of two thousand years. We evaluate the performance of seven different classifiers both in a tenfold test on 83,892 words from the aligned parallel corpus and on a smaller, manually annotated sample of 525 words, measuring both the overall accuracy of each system and how well that accuracy correlates (via mean square error) to the observed historical variation.