SHARED TASKS AND COMPARATIVE EVALUATION IN NATURAL LANGUAGE GENERATION

SHARED TASKS AND COMPARATIVE EVALUATION IN NATURAL LANGUAGE GENERATION
复制标题

自然语言生成中的共享任务和比较评估

DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
D. McDonald
D. McDonald
中科院分区:
--
文献类型:
--
作者:
D. McDonald

文献摘要

被引文献

相似文献

当今的NLG努力应与实际的人类绩效进行比较,这是流利的,并且随机而变化,因此不应针对固定的“黄金标准”文本进行评估,并且共同的任务努力不应假设他们可以贴上表示形式源内容的源头,但玩家仍会产生现实世界所要求的文本的多样性。自然语言产生(NLG)系统是一个人的输出,除了偶尔出现语音错误或其他预测的分散体,例如口吃或重新启动,人们完全指挥自己的语法,并完全指挥他们的话语上下文,因为它塑造了他们所说的一致性以及他们说的任何NLG系统的凝聚力。他们几乎),当他们描述随后对已经引入的话语中引入的实体的引用时,这并不能减少复杂的NP,这不会在连接时与共同主题相结合,或者无法使用其他任何其他普通粘性技术可用他们使用的语言根本不在运行中。
Today’s NLG efforts should be compared against actual human performance, which is fluent and varies randomly and with context. Consequently, evaluations should not be done against a fixed ‘gold standard’ text, and shared task efforts should not assume that they can stipulate the representation of the source content and still let players generate the diversity of texts that the real world calls for. 1 Minimal competency The proper point of reference when making an evaluation of the output of a natural language generation (NLG) system is the output of a person. With the exception of the occasional speech error or other predicable disfluencies such as stuttering or restarts, people speak with complete command of their grammar (not to mention their culturally attuned prosodics), and with complete command of their discourse context as it shapes the coherence of what they say and the cohesion of how they say it. Any NLG system today that does not use pronouns correctly (assuming they use them at all), that does not reduce complex NPs when they describe subsequent references to entities already introduced into the discourse, that does not reduce clauses with common subjects when they are conjoined, or that fails to use any of the other ordinary cohesive techniques available to them in the language they are using is simply not in the running. Human-level fluency is the entrance ticket to any comparative evaluation of NLG systems.