Web Search Clustering and Labeling with Hidden Topics

Web Search Clustering and Labeling with Hidden Topics
复制标题

DOI:
10.1145/1568292.1568295
复制
发表时间:
2009-08
期刊:
ACM Trans. Asian Lang. Inf. Process.
影响因子:
--
通讯作者:
Cam-Tu Nguyen;X. Phan;S. Horiguchi;Thu-Trang Nguyen;Quang-Thuy Ha
Cam-Tu Nguyen;X. Phan;S. Horiguchi;Thu-Trang Nguyen;Quang-Thuy Ha
中科院分区:
其他
文献类型:
--
作者:
Cam-Tu Nguyen;X. Phan;S. Horiguchi;Thu-Trang Nguyen;Quang-Thuy Ha

文献摘要

被引文献

相似文献

Web搜索聚类是一种以更方便浏览的方式重新组织搜索结果(也称为“片段”)的解决方案。有三个关键的要求,这样的检索后聚类系统:(1)聚类算法应该组相似的文档在一起;(2)集群应该标记有描述性短语;(3)聚类系统应该提供高质量的聚类,而无需下载整个网页。本文介绍了一种新的框架聚类Web搜索结果在越南的目标,上述三个问题。主要的动机是,通过丰富的短片段与隐藏的主题,从巨大的资源,在互联网上的文件,它是能够有效地集群和标签这样的片段在一个面向主题的方式,而不涉及整个网页。我们的方法是基于最近成功的主题分析模型,如概率潜在语义分析,或潜在狄利克雷分配。该框架的基本思想是,我们收集一个非常大的外部数据集合,称为“通用数据集”,然后在原始片段和从通用数据集合中发现的丰富隐藏主题集上构建一个聚类系统。这可以被看作是要聚类的片段的更丰富的表示。我们对我们的方法进行了仔细的评估,结果表明我们的方法可以产生令人印象深刻的聚类质量。
Web search clustering is a solution to reorganize search results (also called “snippets”) in a more convenient way for browsing. There are three key requirements for such post-retrieval clustering systems: (1) the clustering algorithm should group similar documents together; (2) clusters should be labeled with descriptive phrases; and (3) the clustering system should provide high-quality clustering without downloading the whole Web page. This article introduces a novel framework for clustering Web search results in Vietnamese which targets the three above issues. The main motivation is that by enriching short snippets with hidden topics from huge resources of documents on the Internet, it is able to cluster and label such snippets effectively in a topic-oriented manner without concerning whole Web pages. Our approach is based on recent successful topic analysis models, such as Probabilistic-Latent Semantic Analysis, or Latent Dirichlet Allocation. The underlying idea of the framework is that we collect a very large external data collection called “universal dataset,” and then build a clustering system on both the original snippets and a rich set of hidden topics discovered from the universal data collection. This can be seen as a richer representation of snippets to be clustered. We carry out careful evaluation of our method and show that our method can yield impressive clustering quality.