VSoLSCSum: Building a Vietnamese Sentence-Comment Dataset for Social Context Summarization

VSoLSCSum: Building a Vietnamese Sentence-Comment Dataset for Social Context Summarization
复制标题

DOI:
--
复制
发表时间:
2016-12
期刊:
--
影响因子:
--
通讯作者:
Minh-Tien Nguyen;Lai Dac Viet;Phong Do;Vu Tran;M. Nguyen
Minh-Tien Nguyen;Lai Dac Viet;Phong Do;Vu Tran;M. Nguyen
中科院分区:
其他
文献类型:
--
作者:
Minh-Tien Nguyen;Lai Dac Viet;Phong Do;Vu Tran;M. Nguyen

文献摘要

相似文献

本文介绍了VSoLSCSum,一个越南语链接句子评论数据集,该数据集是手动创建的,用于解决越南语社会语境总结缺乏标准语料库的问题。该数据集是通过在越南网页上提到的12个特殊事件中的141个Web文档的关键词收集的。社交用户被要求参与创建标准摘要和每个句子或评论的标签。经验证后,由Cohen’s Kappa计算得出的评级者间的协议间为0.685。为了说明我们数据集的潜在用途,通过使用一组局部和社会特征来训练学习排序方法。实验结果表明,在我们的数据集上训练的摘要模型在ROUGE-1和ROUGE-2的社会背景摘要中都优于最先进的基线。
This paper presents VSoLSCSum, a Vietnamese linked sentence-comment dataset, which was manually created to treat the lack of standard corpora for social context summarization in Vietnamese. The dataset was collected through the keywords of 141 Web documents in 12 special events, which were mentioned on Vietnamese Web pages. Social users were asked to involve in creating standard summaries and the label of each sentence or comment. The inter-agreement calculated by Cohen’s Kappa among raters after validating is 0.685. To illustrate the potential use of our dataset, a learning to rank method was trained by using a set of local and social features. Experimental results indicate that the summary model trained on our dataset outperforms state-of-the-art baselines in both ROUGE-1 and ROUGE-2 in social context summarization.