Overview of MSLR2022: A Shared Task on Multi-document Summarization for Literature Reviews

Overview of MSLR2022: A Shared Task on Multi-document Summarization for Literature Reviews
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Lucy Lu Wang;Jay DeYoung;Byron Wallace
Lucy Lu Wang;Jay DeYoung;Byron Wallace
中科院分区:
其他
文献类型:
--
作者:
Lucy Lu Wang;Jay DeYoung;Byron Wallace

文献摘要

相似文献

我们概述了MSLR2022在多文档摘要上的共享任务,用于文献综述。这项共享任务是在COLING 2022第三届学术文件处理(SDP)研讨会上主持的。对于这项任务,我们提供了由从综述论文中提取的金摘要以及合成到这些摘要中的输入摘要组组成的数据,这些摘要分为两个子任务。总共有6个小组参与,提交了10份公开提交,其中6份提交给Cochrane子任务,4份提交给MS²子任务。得分最高的系统在Cochrane子任务上报告了超过2分的ROUGE-L改进,尽管在所有自动评估指标上并没有一致地报告性能改进;对结果的定性检查也表明,目前的评价量度不足以捕捉这一任务的真实性和一致性。需要做大量的工作来改进系统性能,更重要的是,需要开发更好的方法来自动评估该任务的性能。
We provide an overview of the MSLR2022 shared task on multi-document summarization for literature reviews. The shared task was hosted at the Third Scholarly Document Processing (SDP) Workshop at COLING 2022. For this task, we provided data consisting of gold summaries extracted from review papers along with the groups of input abstracts that were synthesized into these summaries, split into two subtasks. In total, six teams participated, making 10 public submissions, 6 to the Cochrane subtask and 4 to the MSˆ2 subtask. The top scoring systems reported over 2 points ROUGE-L improvement on the Cochrane subtask, though performance improvements are not consistently reported across all automated evaluation metrics; qualitative examination of the results also suggests the inadequacy of current evaluation metrics for capturing factuality and consistency on this task. Significant work is needed to improve system performance, and more importantly, to develop better methods for automatically evaluating performance on this task.