The Discussion Tracker Corpus of Collaborative Argumentation

The Discussion Tracker Corpus of Collaborative Argumentation
复制标题

DOI:
--
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Christopher Olshefski;Luca Lugini;Ravneet Singh;D. Litman;Amanda Godley
Christopher Olshefski;Luca Lugini;Ravneet Singh;D. Litman;Amanda Godley
中科院分区:
其他
文献类型:
--
作者:
Christopher Olshefski;Luca Lugini;Ravneet Singh;D. Litman;Amanda Godley

文献摘要

被引文献

相似文献

尽管近年来关于论据挖掘的NLP研究取得了很大进展,但大多数研究都是利用异步和书面文本的语料库,这些语料库通常是由个人产生的。同步、多方论证的公开语料库很少。讨论跟踪语料库收集于高中英语课堂,是一个带注释的口语多方辩论文本数据集。该语料库包括从985分钟的音频中转录的29个多方英国文学讨论。抄本被标注为协作论证的三个维度:论证移动(主张、证据和解释)、特异性(低、中、高)和协作性(例如,对他人观点的扩展和分歧)。除了提供语料库的描述性统计数据外,我们还提供了性能基准和相关代码,分别预测每个维度,说明使用语料库中的多个注释来通过多任务学习提高性能,最后讨论了语料库可能用于进一步NLP研究的其他方法。
Although NLP research on argument mining has advanced considerably in recent years, most studies draw on corpora of asynchronous and written texts, often produced by individuals. Few published corpora of synchronous, multi-party argumentation are available. The Discussion Tracker corpus, collected in high school English classes, is an annotated dataset of transcripts of spoken, multi-party argumentation. The corpus consists of 29 multi-party discussions of English literature transcribed from 985 minutes of audio. The transcripts were annotated for three dimensions of collaborative argumentation: argument moves (claims, evidence, and explanations), specificity (low, medium, high) and collaboration (e.g., extensions of and disagreements about others’ ideas). In addition to providing descriptive statistics on the corpus, we provide performance benchmarks and associated code for predicting each dimension separately, illustrate the use of the multiple annotations in the corpus to improve performance via multi-task learning, and finally discuss other ways the corpus might be used to further NLP research.