Developing a Dataset of Overridden Information in Wikipedia

Developing a Dataset of Overridden Information in Wikipedia
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Masatoshi Tsuchiya;Yasutaka Yokoi
Masatoshi Tsuchiya;Yasutaka Yokoi
中科院分区:
其他
文献类型:
--
作者:
Masatoshi Tsuchiya;Yasutaka Yokoi

文献摘要

相似文献

本文提出了检测信息覆盖的新任务。由于网络上的所有信息都没有及时更新,因此有必要丢弃被其他信息源覆盖的信息。该任务被形式化为二元分类问题,以确定参考句子是否覆盖了目标句子。在研究这项任务时,本文描述了通过从维基百科两个版本之间的差异收集句子对来构建覆盖信息数据集的过程。我们正在开发的数据集表明,旧版本的维基百科包含许多被覆盖的信息,并且信息覆盖的检测是必要的。
This paper proposes a new task of detecting information override. Since all information on the Web is not updated in a timely manner, the necessity is created for information that is overridden by another information source to be discarded. The task is formalized as a binary classification problem to determine whether a reference sentence has overridden a target sentence. In investigating this task, this paper describes a construction procedure for the dataset of overridden information by collecting sentence pairs from the difference between two versions of Wikipedia. Our developing dataset shows that the old version of Wikipedia contains much overridden information and that the detection of information override is necessary.