Recognition of Pornographic Web Pages by Classifying Texts and Images

Recognition of Pornographic Web Pages by Classifying Texts and Images
复制标题

DOI:
10.1109/tpami.2007.1133
复制
发表时间:
2007-06
影响因子:
23.6
通讯作者:
Weiming Hu;Ou Wu;Zhouyao Chen;Zhouyu Fu;S. Maybank
Weiming Hu;Ou Wu;Zhouyao Chen;Zhouyu Fu;S. Maybank
中科院分区:
计算机科学1区
文献类型:
--
作者:
Weiming Hu;Ou Wu;Zhouyao Chen;Zhouyu Fu;S. Maybank

文献摘要

被引文献

相似文献

随着万维网的快速发展,人们从信息共享中受益越来越多。然而,含有淫秽、有害或非法内容的网页却很容易被访问。识别此类不合适、攻击性或色情网页非常重要。在本文中,描述了一种用于识别色情网页的新颖框架。 C4.5决策树用于根据内容表示将网页分为连续文本页面、离散文本页面和图像页面。这三类网页分别由连续文本分类器、离散文本分类器以及融合图像分类器和离散文本分类器结果的算法来处理。在连续文本分类器中,统计和语义特征用于识别色情文本。在离散文本分类器中,使用朴素贝叶斯规则来计算离散文本是色情的概率。在图像分类器中,提取对象的基于轮廓的特征来识别色情图像。在文本和图像融合算法中,利用贝叶斯理论将图像和文本的识别结果结合起来。实验结果表明,连续文本分类器优于传统的基于关键字统计的分类器,基于轮廓的图像分类器优于传统的基于皮肤区域的图像分类器,我们的融合算法获得的结果优于任何一个单独的分类器,并且我们的框架可以适应不同类别的网页
With the rapid development of the World Wide Web, people benefit more and more from the sharing of information. However, Web pages with obscene, harmful, or illegal content can be easily accessed. It is important to recognize such unsuitable, offensive, or pornographic Web pages. In this paper, a novel framework for recognizing pornographic Web pages is described. A C4.5 decision tree is used to divide Web pages, according to content representations, into continuous text pages, discrete text pages, and image pages. These three categories of Web pages are handled, respectively, by a continuous text classifier, a discrete text classifier, and an algorithm that fuses the results from the image classifier and the discrete text classifier. In the continuous text classifier, statistical and semantic features are used to recognize pornographic texts. In the discrete text classifier, the naive Bayes rule is used to calculate the probability that a discrete text is pornographic. In the image classifier, the object's contour-based features are extracted to recognize pornographic images. In the text and image fusion algorithm, the Bayes theory is used to combine the recognition results from images and texts. Experimental results demonstrate that the continuous text classifier outperforms the traditional keyword-statistics-based classifier, the contour-based image classifier outperforms the traditional skin-region-based image classifier, the results obtained by our fusion algorithm outperform those by either of the individual classifiers, and our framework can be adapted to different categories of Web pages