Classifying Wikipedia articles using network motif counts and ratios

Classifying Wikipedia articles using network motif counts and ratios
复制标题

使用网络主题计数和比率对维基百科文章进行分类

DOI:
10.1145/2462932.2462948
复制
发表时间:
2012
期刊:
2017 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM)
影响因子:
--
通讯作者:
P. Cunningham
P. Cunningham
中科院分区:
--
文献类型:
--
作者:
Guangyu Wu;Martin Harrigan;P. Cunningham

文献摘要

被引文献

相似文献

由于维基百科文章的制作是一个协作过程,因此围绕文章的编辑网络可以告诉我们有关该文章质量的信息。没有受到多少关注的文章,网络就会稀疏;另一方面,作为维基百科战场的文章将拥有非常拥挤的网络。在本文中,我们评估了将编辑网络表征为可用于聚类和分类的主题计数向量的想法。我们的目标不是立即开发一个强大的分类器,而是评估网络主题中的信号是什么。我们证明,这种主题计数向量表示对于在维基百科质量等级上对文章进行分类是有效的。我们进一步表明,在比较大小完全不同的网络时,基序计数的比率可以有效克服归一化问题。
Because the production of Wikipedia articles is a collaborative process, the edit network around a article can tell us something about the quality of that article. Articles that have received little attention will have sparse networks; at the other end of the spectrum, articles that are Wikipedia battle grounds will have very crowded networks. In this paper we evaluate the idea of characterizing edit networks as a vector of motif counts that can be used in clustering and classification. Our objective is not immediately to develop a powerful classifier but to assess what is the signal in network motifs. We show that this motif count vector representation is effective for classifying articles on the Wikipedia quality scale. We further show that ratios of motif counts can effectively overcome normalization problems when comparing networks of radically different sizes.