Query-driven document partitioning and collection selection

Query-driven document partitioning and collection selection
复制标题

查询驱动的文档分区和集合选择

DOI:
--
复制
发表时间:
2006
期刊:
Scalable Information Systems
影响因子:
--
通讯作者:
D. Laforenza
D. Laforenza
中科院分区:
--
文献类型:
--
作者:
D. Puppin;F. Silvestri;D. Laforenza

文献摘要

被引文献

相似文献

我们提出了一种新的策略,分区的文档集合到几个服务器上,并进行有效的收集选择。该方法基于对查询日志的分析。我们提出了一种新的文档表示称为查询向量模型。每个文档都被表示为一个列表,该列表记录了文档本身匹配的查询及其排名。为了划分集合并构建集合选择函数,我们将查询和文档进行共聚类。然后将文档聚类分配给底层IR服务器,而查询聚类表示返回类似结果的查询,并用于集合选择。我们表明,这种文档分区策略大大提高了标准的集合选择算法,包括科里,w.r.t.循环分配其次,我们表明,通过匹配查询到现有的查询集群,并连续选择只有一个服务器,执行集合选择,我们达到了平均精度-在-5高达1.74,我们不断提高科里精度的因素之间的11%和15%。作为一个附带的结果,我们展示了一种选择很少要求的文档的方法。将这些文档与集合的其余部分分开,可以使索引器生成一个更紧凑的索引,其中只包含将来可能被请求的相关文档。在我们的测试中,大约52%的文档(3,128,366)没有在任何查询的前100个排名靠前的结果中返回。
We present a novel strategy to partition a document collection onto several servers and to perform effective collection selection. The method is based on the analysis of query logs. We proposed a novel document representation called query-vectors model. Each document is represented as a list recording the queries for which the document itself is a match, along with their ranks. To both partition the collection and build the collection selection function, we co-cluster queries and documents. The document clusters are then assigned to the underlying IR servers, while the query clusters represent queries that return similar results, and are used for collection selection. We show that this document partition strategy greatly boosts the performance of standard collection selection algorithms, including CORI, w.r.t. a round-robin assignment. Secondly, we show that performing collection selection by matching the query to the existing query clusters and successively choosing only one server, we reach an average precision-at-5 up to 1.74 and we constantly improve CORI precision of a factor between 11% and 15%. As a side result we show a way to select rarely asked-for documents. Separating these documents from the rest of the collection allows the indexer to produce a more compact index containing only relevant documents that are likely to be requested in the future. In our tests, around 52% of the documents (3,128,366) are not returned among the first 100 top-ranked results of any query.