Detecting Censorable Content on Sina Weibo: A Pilot Study

Detecting Censorable Content on Sina Weibo: A Pilot Study
复制标题

DOI:
10.1145/3200947.3201037
复制
发表时间:
2018-07
期刊:
Proceedings of the 10th Hellenic Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Kei Yin Ng;Anna Feldman;C. Leberknight
Kei Yin Ng;Anna Feldman;C. Leberknight
中科院分区:
其他
文献类型:
--
作者:
Kei Yin Ng;Anna Feldman;C. Leberknight

文献摘要

被引文献

相似文献

本研究对中国大陆网络审查的语言特征进行了初步的探讨。我们收集了344个新浪微博上发布的经过审查和未经审查的微博帖子的语料库,并基于语言,主题无关的特征构建了一个朴素贝叶斯分类器。该分类器在预测新浪微博上的博客文章是否会被审查方面达到了79.34%的准确率。
This study provides preliminary insights into the linguistic features that contribute to Internet censorship in mainland China. We collected a corpus of 344 censored and uncensored microblog posts that were published on Sina Weibo and built a Naive Bayes classifier based on the linguistic, topic-independent, features. The classifier achieves a 79.34% accuracy in predicting whether a blog post would be censored on Sina Weibo.