DebateSum: A large-scale argument mining and summarization dataset

DebateSum: A large-scale argument mining and summarization dataset
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Allen Roush;Arvind Balaji
Allen Roush;Arvind Balaji
中科院分区:
其他
文献类型:
--
作者:
Allen Roush;Arvind Balaji

文献摘要

被引文献

相似文献

先前在论证挖掘方面的工作经常提到它在自动辩论系统中的潜在应用。尽管如此,几乎没有数据集或模型将自然语言处理技术应用于竞争性正式辩论中发现的问题。为了解决这个问题,我们提出了DebateSum数据集。DebateSum包含187,386个独特的证据,以及相应的论点和摘要。DebateSum是使用国家演讲和辩论协会的竞争对手在7年内收集的数据制作的。我们训练了几个Transformer摘要模型,以在DebateSum上测试摘要性能。我们还介绍了一组在DebateSum上训练的快速文本词向量,称为debate2vec。最后,我们提出了一个搜索引擎,这个数据集被广泛使用的国家演讲和辩论协会的成员今天。DebateSum搜索引擎可供公众访问:http://www.debate.cards
Prior work in Argument Mining frequently alludes to its potential applications in automatic debating systems. Despite this focus, almost no datasets or models exist which apply natural language processing techniques to problems found within competitive formal debate. To remedy this, we present the DebateSum dataset. DebateSum consists of 187,386 unique pieces of evidence with corresponding argument and extractive summaries. DebateSum was made using data compiled by competitors within the National Speech and Debate Association over a 7year period. We train several transformer summarization models to benchmark summarization performance on DebateSum. We also introduce a set of fasttext word-vectors trained on DebateSum called debate2vec. Finally, we present a search engine for this dataset which is utilized extensively by members of the National Speech and Debate Association today. The DebateSum search engine is available to the public here: http://www.debate.cards