Controversy Detection in Wikipedia Using Collective Classification
Controversy Detection in Wikipedia Using Collective Classification
复制标题
使用集体分类检测维基百科中的争议
DOI:
10.1145/2911451.2914745
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
J. Allan
中科院分区:
文献类型:
--
作者:
Shiri Dori;David D. Jensen;J. Allan
Concerns over personalization in IR have sparked an interest in detection and analysis of controversial topics. Accurate detection would enable many beneficial applications, such as alerting search users to controversy. Wikipedia's broad coverage and rich metadata offer a valuable resource for this problem. We hypothesize that intensities of controversy among related pages are not independent; thus, we propose a stacked model which exploits the dependencies among related pages. Our approach improves classification of controversial web pages when compared to a model that examines each page in isolation, demonstrating that controversial topics exhibit homophily. Using notions of similarity to construct a subnetwork for collective classification, rather than using the default network present in the relational data, leads to improved classification with wider applications for semi-structured datasets, with the effects most pronounced when a small set of neighbors is used.