Combining link-based and content-based methods for web document classification

Combining link-based and content-based methods for web document classification
复制标题

DOI:
10.1145/956863.956938
复制
发表时间:
2003-11
期刊:
--
影响因子:
--
通讯作者:
P. Calado;Marco Cristo;E. Moura;N. Ziviani;B. Ribeiro-Neto;Marcos André Gonçalves
P. Calado;Marco Cristo;E. Moura;N. Ziviani;B. Ribeiro-Neto;Marcos André Gonçalves
中科院分区:
其他
文献类型:
--
作者:
P. Calado;Marco Cristo;E. Moura;N. Ziviani;B. Ribeiro-Neto;Marcos André Gonçalves

文献摘要

被引文献

相似文献

本文研究如何链接信息可以用来改善分类结果的Web收藏。我们评估四种不同的措施,来自Web链接结构的主题相似性,并确定它们在预测文档类别的准确性。使用贝叶斯网络模型,我们结合联合收割机这些措施与传统的基于内容的分类器所获得的结果。在Web目录上的实验表明,当考虑目录外页面的链接时,可以获得最佳结果。与传统的基于内容的分类器相比,仅链接信息就能够在F1中获得高达46点的增益。与基于内容的方法相结合可以进一步改善结果,但可能会引入太多的噪音,因为Web页面的文本是一个不太可靠的信息源。这项工作提供了一个重要的洞察力,从链接的措施更适合比较Web文档,以及如何将这些措施与基于内容的算法相结合,以提高Web分类的有效性。
This paper studies how link information can be used to improve classification results for Web collections. We evaluate four different measures of subject similarity, derived from the Web link structure, and determine how accurate they are in predicting document categories. Using a Bayesian network model, we combine these measures with the results obtained by traditional content-based classifiers. Experiments on a Web directory show that best results are achieved when links from pages outside the directory are considered. Link information alone is able to obtain gains of up to 46 points in F1, when compared to a traditional content-based classifier. The combination with content-based methods can further improve the results, but too much noise may be introduced, since the text of Web pages is a much less reliable source of information. This work provides an important insight on which measures derived from links are more appropriate to compare Web documents and how these measures can be combined with content-based algorithms to improve the effectiveness of Web classification.