Controversy Detection in Wikipedia Using Collective Classification

Controversy Detection in Wikipedia Using Collective Classification
复制标题

使用集体分类检测维基百科中的争议

DOI:
10.1145/2911451.2914745
复制
发表时间:
2016
期刊:
Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
J. Allan
J. Allan
中科院分区:
--
文献类型:
--
作者:
Shiri Dori;David D. Jensen;J. Allan

文献摘要

被引文献

相似文献

对IR中个性化的关注引发了对有争议的主题的检测和分析的兴趣。准确的检测将使许多有益的应用,如提醒搜索用户的争议。维基百科的广泛覆盖和丰富的元数据为这个问题提供了宝贵的资源。我们假设,相关页面之间的争议强度是不独立的,因此,我们提出了一个堆栈模型,利用相关页面之间的依赖关系。我们的方法提高了分类有争议的网页相比,孤立地检查每个页面的模型,表明有争议的主题表现出同质性。使用相似性的概念来构造用于集体分类的子网络,而不是使用关系数据中存在的默认网络,导致改进的分类,具有半结构化数据集的更广泛应用,当使用一小组邻居时效果最明显。
Concerns over personalization in IR have sparked an interest in detection and analysis of controversial topics. Accurate detection would enable many beneficial applications, such as alerting search users to controversy. Wikipedia's broad coverage and rich metadata offer a valuable resource for this problem. We hypothesize that intensities of controversy among related pages are not independent; thus, we propose a stacked model which exploits the dependencies among related pages. Our approach improves classification of controversial web pages when compared to a model that examines each page in isolation, demonstrating that controversial topics exhibit homophily. Using notions of similarity to construct a subnetwork for collective classification, rather than using the default network present in the relational data, leads to improved classification with wider applications for semi-structured datasets, with the effects most pronounced when a small set of neighbors is used.