A Tolerance Rough Set Approach to Clustering Web Search Results

A Tolerance Rough Set Approach to Clustering Web Search Results
复制标题

DOI:
10.1007/978-3-540-30116-5_51
复制
发表时间:
2004-09
期刊:
2012 8th International Conference on Informatics and Systems (INFOS)
影响因子:
--
通讯作者:
Chi Lang Ngo;H. Nguyen
Chi Lang Ngo;H. Nguyen
中科院分区:
其他
文献类型:
--
作者:
Chi Lang Ngo;H. Nguyen

文献摘要

被引文献

相似文献

两个最流行的方法,以促进在网络上搜索信息是由网络搜索引擎和网络目录。虽然搜索引擎的性能每天都在提高,但由于Web的巨大规模和高度动态性,在Web上搜索可能是一项繁琐而耗时的任务。此外,用户的“搜索背后的意图”没有明确表达,这导致过于笼统,简短的查询。搜索引擎返回的结果可以从几百到几十万个文档,聚类是一种管理大量结果的方法。搜索结果聚类可以定义阿萨自动将搜索结果分组到主题组的过程。然而,与传统的文档聚类相比,搜索结果的聚类是在运行中(每个用户查询请求)和本地从搜索引擎返回的有限结果集上完成的。搜索结果的聚类可以帮助用户更有效地浏览大量文档。本文在文献聚类研究的基础上,提出了一种基于容差粗糙集的搜索结果聚类方法。公差等级用于近似文档中存在的概念。将容差粗糙集模型应用于文本聚类,丰富了文本和聚类的表示形式,提高了聚类性能。
Two most popular approaches to facilitate searching for information on the web are represented by web search engine and web directories. Although the performance of search engines is improving every day, searching on the web can be a tedious and time-consuming task due to the huge size and highly dynamic nature of the web. Moreover, the user’s “intention behind the search” is not clearly expressed which results in too general, short queries. Results returned by search engine can count from hundreds to hundreds of thousands of documents.One approach to manage the large number of results is clustering. Search results clustering can be defined asa process of automatical grouping search results into to thematic groups. However, in contrast to traditional document clustering, clustering of search results are done on-the-fly (per user query request) and locally on a limited set of results return from the search engine. Clustering of search results can help user navigate through large set of documents more efficiently. By providing concise, accurate description of clusters, it lets user localizes interesting document faster.In this paper, we proposed an approach to search results clustering based on Tolerance Rough Set following the work on document clustering [4,3]. Tolerance classes are used to approximate concepts existed in documents. The application of Tolerance Rough Set model in document clustering was proposed as a way to enrich document and cluster representation with the hope of increasing clustering performance.