Maximum Entropy Models and Stochastic Optimality Theory

Maximum Entropy Models and Stochastic Optimality Theory
复制标题

最大熵模型和随机最优理论

DOI:
--
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
Gerhard Jäger
Gerhard Jäger
中科院分区:
--
文献类型:
--
作者:
Gerhard Jäger

文献摘要

被引文献

相似文献

在最近的一系列出版物中(最著名的是Boersma(1998);另见Boersma和Hayes(2001)),Paul Boersma在Prince和Smolensky(1993)的意义上发展了标准最优性理论的随机推广。虽然经典的OT文法将一组候选映射到其最优元素(或多个元素),但在Boersma的随机最优理论(StOT)中,文法定义了这样一个集合上的概率分布。Boersma还开发了一种自然学习算法,渐进学习算法(GLA),它从语料库中归纳出StOT语法。StOT能够科普自然语言现象,如歧义,可选性和梯度语法性,这些对于标准OT来说是出了名的问题。Keller和Asudeh(2002)对StOT提出了一些批评,特别是GLA。Goldwater和约翰逊(2003)指出,在计算语言学中广泛使用的最大熵(ME)模型可能是StOT的替代品。ME模型与StOT足够相似,使得在前一个模型中获得的经验结果可以转移到后者,并且这些模型可以说比StOT具有更好的形式属性。另一方面,GLA比ME模型的标准学习算法具有更高的认知可扩展性(从Boersma和Levelt(2000)可以看出)。在本文中,我将论证联合收割机与ME模型的优点结合是可能的。它可以表明,GLA可以适应ME模型几乎没有修改。换句话说,GLA是ME模型的唯一最自然的在线学习算法。Keller和Asudeh的批评,在一定程度上是合理的,并不适用于ME评估与GLA学习的结合,GLA的认知优势得到了维护。
In a series of recent publications (most notably Boersma (1998); see also Boersma and Hayes (2001)), Paul Boersma has developed a stochastic generalization of standard Optimality Theory in the sense of Prince and Smolensky (1993). While a classical OT grammar maps a set of candidates to its optimal element (or elements), in Boersma’s Stochastic Optimality Theory (StOT for short) a grammar defines a probability distribution over such a set. Boersma also developed a natural learning algorithm, the Gradual Learning Algorithm (GLA) that induces a StOT grammar from a corpus. StOT is able to cope with natural language phenomena like ambiguity, optionality, and gradient grammaticality, that are notoriously problematic for standard OT. Keller and Asudeh (2002) raise several criticisms against StOT in general and the GLA in particular. Partially as a reaction to that, Goldwater and Johnson (2003) point out that maximum entropy (ME) models, that are widely used in computational linguistics, might be an alternative to StOT. ME models are similar enough to StOT to make it possible that empirical results reached in the former model can be transferred to the latter, and these models have arguably better formal properties than StOT. On the other hand, the GLA has a higher cognitive plausibility (as can be seen from Boersma and Levelt (2000)) than the standard learning algorithms for ME models. In this paper I will argue that it is possible to combine the advantages of StOT with the ME model. It can be shown that the GLA can be adapted to ME models almost without modifications. Put differently, it turns out that the GLA is the single most natural on-line learning algorithm for ME models. Keller and Asudeh’s criticism, to the degree that it is justified, does not apply to the combination of ME evaluation with GLA learning, and the cognitive advantages of the GLA are maintained.