Automated Controversy Detection on the Web

Automated Controversy Detection on the Web
复制标题

网络上的自动争议检测

DOI:
--
复制
发表时间:
2015
期刊:
European Conference on Information Retrieval
影响因子:
--
通讯作者:
J. Allan
J. Allan
中科院分区:
--
文献类型:
--
作者:
Shiri Dori;J. Allan

文献摘要

被引文献

相似文献

提醒用户注意有争议的搜索结果可以鼓励批判性阅读,促进健康的公民话语,并抵消“过滤泡沫”效应,因此这将是搜索引擎或浏览器扩展中的一个有用功能。然而,为了实现这样的特征,必须解决确定哪些主题或网页是有争议的二进制分类任务。早期的工作描述了使用受监督的最近邻分类器的概念证明,该分类器可以访问人工注释的维基百科文章的先知。本文通过在弱监督分类方法中利用维基百科文章中丰富的元数据,将人类排除在循环之外,从而概括和扩展了这一概念。我们提出的新技术允许最近邻方法在更大的范围内扩展到其他数据集。结果大大超过了朴素的基线,以F1、F0.5和准确度的标准衡量,几乎与依赖先知的方法相同。最后,我们讨论了解决这一问题的含义,作为IR社区感兴趣的更广泛主题的一部分,并提出了在这一令人兴奋的新领域进一步探索的几种途径。
Alerting users about controversial search results can encourage critical literacy, promote healthy civic discourse and counteract the “filter bubble” effect, and therefore would be a useful feature in a search engine or browser extension. In order to implement such a feature, however, the binary classification task of determining which topics or webpages are controversial must be solved. Earlier work described a proof of concept using a supervised nearest neighbor classifier with access to an oracle of manually annotated Wikipedia articles. This paper generalizes and extends that concept by taking the human out of the loop, leveraging the rich metadata available in Wikipedia articles in a weakly-supervised classification approach. The new technique we present allows the nearest neighbor approach to be extended on a much larger scale and to other datasets. The results improve substantially over naive baselines and are nearly identical to the oracle-reliant approach by standard measures of F 1, F 0.5, and accuracy. Finally, we discuss implications of solving this problem as part of a broader subject of interest to the IR community, and suggest several avenues for further exploration in this exciting new space.