Detection of Mergeable Wikipedia Articles Utilizing Multiple Similarity Measures

Detection of Mergeable Wikipedia Articles Utilizing Multiple Similarity Measures
复制标题

DOI:
10.2197/ipsjjip.28.178
复制
发表时间:
2020-01
期刊:
J. Inf. Process.
影响因子:
--
通讯作者:
Renzhi Wang;M. Iwaihara
Renzhi Wang;M. Iwaihara
中科院分区:
其他
文献类型:
--
作者:
Renzhi Wang;M. Iwaihara

文献摘要

相似文献

维基百科是最大的在线百科全书,其中的文章由不同的志愿者以不同的思想和风格编辑。有时两篇或多篇文章的标题不同,但这些文章的主题完全相同或非常相似。管理员和编辑应该检测这样的文章对,并确定它们是否应该合并在一起。我们称一个文章对是可合并的,如果它被讨论为可能的合并,合并的文章对是这样的,该对实际上是合并的。在本文中,我们提出了一种方法来自动确定是否文章对是可合并或合并。根据维基百科条目合并指南,在重复的情况下,条目对覆盖完全相同的内容。在重叠的情况下,文章对覆盖具有显著重叠的相关主题。重叠部分的内容是相似的,但成对的词可能会有很大的不同,所以利用语义相关性的方法是必要的。我们考虑各种文本的相似性和语义相关性。为了整合目标数据集和全球大型语料库上的词嵌入,我们提出了多个嵌入结果的线性和非线性组合,并重建词向量以评估语义相关性。我们澄清了我们的方法和以前的研究之间的差异,结合多个词嵌入。我们还通过计算文章对之间的Jaccard相似度来处理重叠情况。我们结合联合收割机Jaccard相似度,共同链接的文章计数和基于词嵌入的相关性在一起,预测文章对是否应该合并。我们探讨了段级(段落级)相似性与可合并/已合并文章对之间的关系,提出了基于多模态相似性的合并预测(MSBMP),该预测结合了随机森林提出的新特征,用于预测可合并/已合并文章对。我们的评估是在真实的可合并和合并的文章对上进行的。MSBMP的显着优势,从WikiSearch,TFIDF和词嵌入的基线有明显的改善。
Wikipedia is the largest online encyclopedia, in which articles are edited by different volunteers with different thoughts and styles. Sometimes two or more articles’ titles are different but the themes of these articles are exactly the same or strongly similar. Administrators and editors are supposed to detect such article pairs and determine whether they should be merged together. We call an article pair is mergeable if it is discussed for possible merge, and a merged article pair is such that the pair is actually merged. In this paper, we propose a method to automatically determine whether an article pair is mergeable or merged. According to Wikipedia Guidelines for article merge, in the duplicate case, the article pairs are covering exactly the same contents. In the overlap case, the article pairs are covering related subjects that have a significant overlap. The content of an overlapped part is similar but the words in the pair can be extensively different, so methods that exploit semantic relatedness are necessary. We consider various textual similarities and semantic relatedness. For integrating word embeddings on the target dataset and the global large corpus, we propose linear and non-linear combinations of multiple embedding results and rebuilding word vectors for evaluating semantic relatedness. We clarify the differences between our method and previous researches for combining multiple word embeddings. We also deal with overlap cases by computing Jaccard similarity between article pairs. We combine Jaccard similarity, common-link article count and word embedding-based relatedness together, to predict whether the article pair should be merged. We explore the relationship between segment-level (paragraph-level) similarity and mergeable/merged article pairs, then propose Multimodal Similarity-Based Merge Prediction (MSBMP) which combines the proposed new features by Random Forest, to predict mergeable/merged article pairs. Our evaluations are performed on real mergeable and merged article pairs. Remarkable superiorities of MSBMP are shown, with apparent improvement from baselines of WikiSearch, TFIDF and word embeddings.