Automatic Generation of Review Matrices as Multi-document Summarization of Scientific Papers

Automatic Generation of Review Matrices as Multi-document Summarization of Scientific Papers
复制标题

DOI:
--
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
Hayato Hashimoto;Kazutoshi Shinoda;Hikaru Yokono;Akiko Aizawa
Hayato Hashimoto;Kazutoshi Shinoda;Hikaru Yokono;Akiko Aizawa
中科院分区:
其他
文献类型:
--
作者:
Hayato Hashimoto;Kazutoshi Shinoda;Hikaru Yokono;Akiko Aizawa

文献摘要

相似文献

综合矩阵是总结多个文档的各个方面的表格。在我们的工作中,我们专门研究了一个问题,自动生成一个综合矩阵的科学文献综述。如本文所述,我们首先制定的任务作为多文档摘要和问答任务给出了一组方面的审查的基础上调查的系统摘要表的自然语言处理任务。接下来,我们提出了一种方法来解决前一种类型的任务。我们的系统包括两个步骤:句子排序和句子选择。在句子排序步骤中,系统通过将方面视为查询来对输入论文中的句子进行排序。我们使用LexRank,并将查询扩展和单词嵌入,以弥补简洁表达的查询。在句子选择步骤中,系统选择保留在最终输出中的句子。特别强调的摘要类型方面,我们认为这一步作为一个整数线性规划问题,施加一个特殊类型的约束,使摘要可比。我们使用从ACL Anthology创建的数据集评估了我们的系统。人工评价的结果表明,我们的选择方法,利用可比性改善
A synthesis matrix is a table that summarizes various aspects of multiple documents. In our work, we specifically examine a problem of automatically generating a synthesis matrix for scientific literature review. As described in this paper, we first formulate the task as multidocument summarization and question-answering tasks given a set of aspects of the review based on an investigation of system summary tables of NLP tasks. Next, we present a method to address the former type of task. Our system consists of two steps: sentence ranking and sentence selection. In the sentence ranking step, the system ranks sentences in the input papers by regarding aspects as queries. We use LexRank and also incorporate query expansion and word embedding to compensate for tersely expressed queries. In the sentence selection step, the system selects sentences that remain in the final output. Specifically emphasizing the summarization type aspects, we regard this step as an integer linear programming problem with a special type of constraint imposed to make summaries comparable. We evaluated our system using a dataset we created from the ACL Anthology. The results of manual evaluation demonstrated that our selection method using comparability improved