StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning

StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning
复制标题

DOI:
10.48550/arxiv.2210.08459
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Hong Chen;Duc Minh Vo;Hiroya Takamura;Yusuke Miyao;Hideki Nakayama
Hong Chen;Duc Minh Vo;Hiroya Takamura;Yusuke Miyao;Hideki Nakayama
中科院分区:
其他
文献类型:
--
作者:
Hong Chen;Duc Minh Vo;Hiroya Takamura;Yusuke Miyao;Hideki Nakayama

文献摘要

相似文献

现有的自动故事评价方法重视故事词汇层次的连贯性,偏离了人类的偏好。我们超越了这一限制,考虑了一种新颖的故事评价方法,它模仿人类在判断故事时的偏好,即StoryER,它由三个子任务组成:排名、评级和推理。无论是机器生成的故事还是人类编写的故事,StoryER要求机器输出1)与人类偏好相对应的偏好分数,2)特定评级及其相应的置信度,以及3)各个方面的评论(例如,为了支持这些任务,我们引入了一个注释良好的数据集,包括(i)100 k个排名的故事对;以及(ii)一组46 k的评分和对故事各个方面的评论。我们在收集的数据集上微调Longformer-Encoder-Decoder(LED),编码器负责偏好评分和方面预测,解码器负责评论生成。我们的综合实验结果表明,每个任务的竞争基准,显示出与人类偏好的高度相关性。此外,我们已经见证了偏好分数,方面评级和评论的联合学习为每个任务带来的收益。我们的数据集和基准可以公开使用,以推进故事评估任务的研究。
Existing automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference.We go beyond this limitation by considering a novel Story Evaluation method that mimics human preference when judging a story, namely StoryER, which consists of three sub-tasks: Ranking, Rating and Reasoning.Given either a machine-generated or a human-written story, StoryER requires the machine to output 1) a preference score that corresponds to human preference, 2) specific ratings and their corresponding confidences and 3) comments for various aspects (e.g., opening, character-shaping).To support these tasks, we introduce a well-annotated dataset comprising (i) 100k ranked story pairs; and (ii) a set of 46k ratings and comments on various aspects of the story.We finetune Longformer-Encoder-Decoder (LED) on the collected dataset, with the encoder responsible for preference score and aspect prediction and the decoder for comment generation.Our comprehensive experiments result a competitive benchmark for each task, showing the high correlation to human preference.In addition, we have witnessed the joint learning of the preference scores, the aspect ratings, and the comments brings gain each single task.Our dataset and benchmarks are publicly available to advance the research of story evaluation tasks.