SHARED TASKS AND COMPARATIVE EVALUATION IN NATURAL LANGUAGE GENERATION
SHARED TASKS AND COMPARATIVE EVALUATION IN NATURAL LANGUAGE GENERATION
复制标题
自然语言生成中的共享任务和比较评估
DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
D. McDonald
中科院分区:
文献类型:
--
作者:
D. McDonald
Today’s NLG efforts should be compared against actual human performance, which is fluent and varies randomly and with context. Consequently, evaluations should not be done against a fixed ‘gold standard’ text, and shared task efforts should not assume that they can stipulate the representation of the source content and still let players generate the diversity of texts that the real world calls for. 1 Minimal competency The proper point of reference when making an evaluation of the output of a natural language generation (NLG) system is the output of a person. With the exception of the occasional speech error or other predicable disfluencies such as stuttering or restarts, people speak with complete command of their grammar (not to mention their culturally attuned prosodics), and with complete command of their discourse context as it shapes the coherence of what they say and the cohesion of how they say it. Any NLG system today that does not use pronouns correctly (assuming they use them at all), that does not reduce complex NPs when they describe subsequent references to entities already introduced into the discourse, that does not reduce clauses with common subjects when they are conjoined, or that fails to use any of the other ordinary cohesive techniques available to them in the language they are using is simply not in the running. Human-level fluency is the entrance ticket to any comparative evaluation of NLG systems.