Discovering Morphological Paradigms from Plain Text Using a Dirichlet Process Mixture Model

Discovering Morphological Paradigms from Plain Text Using a Dirichlet Process Mixture Model
复制标题

使用狄利克雷过程混合模型从纯文本中发现形态范式

DOI:
--
复制
发表时间:
2011
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Jason Eisner
Jason Eisner
中科院分区:
--
文献类型:
--
作者:
Markus Dreyer;Jason Eisner

文献摘要

被引文献

相似文献

我们提出了一种推理算法,将观察到的单词(标记)组织成结构化的屈折范式(类型)。它还自然地预测这些范式中缺少的未观察到的形式的拼写,并发现泛化到完全未观察到的单词的屈折原则(语法)。 我们的数据贝叶斯生成模型明确表示标记、类型、变形、范式和本地条件字符串编辑。它假设屈折词标记是由屈折范式(字符串元组)的无限混合生成的。每个范式都是从图形模型中一次性采样的,其潜在函数是带有要学习的特定于语言的参数的加权有限状态转换器。这些假设自然会导致一种优雅的经验贝叶斯推理过程,该过程利用了蒙特卡罗电磁法、置信传播和动态规划。给定 50--100 个种子范式,添加 1000 万单词的语料库可将词形变形的预测误差降低高达 10%。
We present an inference algorithm that organizes observed words (tokens) into structured inflectional paradigms (types). It also naturally predicts the spelling of unobserved forms that are missing from these paradigms, and discovers inflectional principles (grammar) that generalize to wholly unobserved words. Our Bayesian generative model of the data explicitly represents tokens, types, inflections, paradigms, and locally conditioned string edits. It assumes that inflected word tokens are generated from an infinite mixture of inflectional paradigms (string tuples). Each paradigm is sampled all at once from a graphical model, whose potential functions are weighted finite-state transducers with language-specific parameters to be learned. These assumptions naturally lead to an elegant empirical Bayes inference procedure that exploits Monte Carlo EM, belief propagation, and dynamic programming. Given 50--100 seed paradigms, adding a 10-million-word corpus reduces prediction error for morphological inflections by up to 10%.