Combining link-based and content-based methods for web document classification
Combining link-based and content-based methods for web document classification
复制标题
DOI:
10.1145/956863.956938
复制
发表时间:
2003-11
期刊:
影响因子:
--
通讯作者:
P. Calado;Marco Cristo;E. Moura;N. Ziviani;B. Ribeiro-Neto;Marcos André Gonçalves
中科院分区:
文献类型:
--
作者:
P. Calado;Marco Cristo;E. Moura;N. Ziviani;B. Ribeiro-Neto;Marcos André Gonçalves
This paper studies how link information can be used to improve classification results for Web collections. We evaluate four different measures of subject similarity, derived from the Web link structure, and determine how accurate they are in predicting document categories. Using a Bayesian network model, we combine these measures with the results obtained by traditional content-based classifiers. Experiments on a Web directory show that best results are achieved when links from pages outside the directory are considered. Link information alone is able to obtain gains of up to 46 points in F1, when compared to a traditional content-based classifier. The combination with content-based methods can further improve the results, but too much noise may be introduced, since the text of Web pages is a much less reliable source of information. This work provides an important insight on which measures derived from links are more appropriate to compare Web documents and how these measures can be combined with content-based algorithms to improve the effectiveness of Web classification.