A Topic-Specific Web Crawler with Concept Similarity Context Graph Based on FCA

A Topic-Specific Web Crawler with Concept Similarity Context Graph Based on FCA
复制标题

DOI:
10.1007/978-3-540-85984-0_101
复制
发表时间:
2008-09
期刊:
--
影响因子:
--
通讯作者:
Yuekui Yang;Yajun Du;Jingyu Sun;Yufeng Hai
Yuekui Yang;Yajun Du;Jingyu Sun;Yufeng Hai
中科院分区:
其他
文献类型:
--
作者:
Yuekui Yang;Yajun Du;Jingyu Sun;Yufeng Hai

文献摘要

被引文献

相似文献

随着互联网的飞速发展,主题爬虫在Web数据挖掘中越来越受到人们的重视。深入研究了如何对未访问网址进行排序,提出了概念相似度上下文图的概念,并提出了一种新颖的特定主题网络爬虫方法,该方法通过形式概念分析(FCA)中概念的相似度来计算未访问网址的预测分数,同时提高检索的精确度和召回率。该方法首先利用用户访问过的页面构建概念格,从概念格中提取反映用户查询主题的核心概念,然后根据核心概念与其他概念之间的语义相似度构建概念相似度上下文图。
With Internet growing exponentially, topic-specific web crawler is becoming more and more popular in the web data mining. How to order the unvisited URLs was studied deeply, we present the notion of concept similarity context graph, and propose a novel approach to topic-specific web crawler, which calculates the unvisited URLs’ prediction score by concepts’ similarity in Formal Concept Analysis (FCA), while improving the retrieval precision and recall ratio. We firstly build a concept lattice using the visited pages, extract the core concepts which reflect the user’s query topic from the concept lattice, and then construct our concept similarity context graph based on the semantic similarities between the core concepts and other concepts.