Query-based sampling of text databases

Query-based sampling of text databases
复制标题

DOI:
10.1145/382979.383040
复制
发表时间:
2001-04
期刊:
ACM Trans. Inf. Syst.
影响因子:
--
通讯作者:
Jamie Callan;Margaret E. Connell
Jamie Callan;Margaret E. Connell
中科院分区:
其他
文献类型:
--
作者:
Jamie Callan;Margaret E. Connell

文献摘要

被引文献

相似文献

企业网络和互联网上的可搜索文本数据库的激增给许多人带来了数据库选择的问题。GGLOSS和CORI等算法可以自动选择搜索给定信息需求的文本数据库,但前提是给出一组准确代表每个数据库内容的资源描述。现有的获取资源描述的技术在用于多方控制的广域网络时具有很大的局限性。本文提出了一种获取准确资源描述的新技术--基于查询的抽样技术。基于查询的抽样不需要资源提供者的合作,也不要求资源提供者使用特定的搜索引擎或表示技术。一组广泛的实验结果表明,资源描述是准确的,计算和通信成本是合理的,并且资源描述确实能够实现准确的自动数据库选择。
The proliferation of searchable text databases on corporate networks and the Internet causes a database selection problem for many people. Algorithms such as gGLOSS and CORI can automatically select which text databases to search for a given information need, but only if given a set of resource descriptions that accurately represent the contents of each database. The existing techniques for a acquiring resource descriptions have significant limitations when used in wide-area networks controlled by many parties. This paper presents query-based sampling, a new technicque for acquiring accurate resource descriptions. Query-based sampling does not require the cooperation of resource providers, nor does it require that resource providers use a particular search engine or representation technique. An extensive set of experimental results demonstrates that accurate resource descriptions are crated, that computation and communication costs are reasonable, and that the resource descriptions do in fact enable accurate automatic dtabase selection.