Phrase-Based Statistical Language Generation Using Graphical Models and Active Learning

Phrase-Based Statistical Language Generation Using Graphical Models and Active Learning
复制标题

DOI:
--
复制
发表时间:
2010-07
期刊:
--
影响因子:
--
通讯作者:
François Mairesse;Milica Gasic;Filip Jurcícek;Simon Keizer;Blaise Thomson;Kai Yu;S. Young
François Mairesse;Milica Gasic;Filip Jurcícek;Simon Keizer;Blaise Thomson;Kai Yu;S. Young
中科院分区:
其他
文献类型:
--
作者:
François Mairesse;Milica Gasic;Filip Jurcícek;Simon Keizer;Blaise Thomson;Kai Yu;S. Young

文献摘要

被引文献

相似文献

先前关于可训练语言生成的大多数工作都集中在两种范式:(a)使用统计模型对一组生成的话语进行排序,或(b)使用统计数据来通知生成决策过程。这两种方法都依赖于手工生成器的存在,这限制了它们对新领域的可扩展性。本文介绍了 Bagel,一种统计语言生成器,它使用动态贝叶斯网络从 42 个未经训练的注释器生成的语义对齐数据中进行学习。人类评估表明,Bagel 可以根据信息呈现领域中看不见的输入生成自然且信息丰富的话语。此外,通过使用基于确定性的主动学习,稀疏数据集的生成性能得到显着提高,用一小部分数据产生接近人类黄金标准的评级。
Most previous work on trainable language generation has focused on two paradigms: (a) using a statistical model to rank a set of generated utterances, or (b) using statistics to inform the generation decision process. Both approaches rely on the existence of a handcrafted generator, which limits their scalability to new domains. This paper presents Bagel, a statistical language generator which uses dynamic Bayesian networks to learn from semantically-aligned data produced by 42 untrained annotators. A human evaluation shows that Bagel can generate natural and informative utterances from unseen inputs in the information presentation domain. Additionally, generation performance on sparse datasets is improved significantly by using certainty-based active learning, yielding ratings close to the human gold standard with a fraction of the data.