Metadata harvesting for content-based distributed information retrieval

Metadata harvesting for content-based distributed information retrieval
复制标题

用于基于内容的分布式信息检索的元数据收集

DOI:
10.1002/asi.20694
复制
发表时间:
2007
影响因子:
--
通讯作者:
Simeoni F
Simeoni F
中科院分区:
--
文献类型:
--
作者:
Simeoni F

文献摘要

参考文献

被引文献

相似文献

我们提出了一种基于内容的分布式信息检索方法,该方法基于广泛分散和自主管理的文档源的全内容索引的定期和增量集中。受开放档案计划(OAI)元数据收集协议的成功启发,该方法占据了内容爬行和分布式检索之间的中间地带。与爬行一样,一些数据会移动到检索过程中,但它是关于内容的统计数据,而不是内容本身;这赠款更有效地利用网络资源和更广泛的应用范围。在分布式检索中,某些处理是随数据沿着分布的,但它是索引而不是检索;这降低了内容提供的成本,同时提高了检索的简单性、有效性和响应性。总的来说,我们认为这种方法保留了集中式检索的良好特性,而不会放弃成本效益高的大规模资源池。我们讨论了与该方法相关的需求,并确定了在OAI基础设施之上部署该方法的两种策略。特别是,我们定义了OAI协议的最小扩展,支持内容资源的完整内容索引和描述性元数据的协调收获。最后,我们报告了一个概念验证原型服务的实现,用于分布式文件集合的多模型基于内容的检索。
We propose an approach to content‐based Distributed Information Retrieval based on the periodic and incremental centralization of full‐content indices of widely dispersed and autonomously managed document sources. Inspired by the success of the Open Archive Initiative's (OAI) Protocol formetadata harvesting, the approach occupies middle ground between content crawling and distributed retrieval. As in crawling, some data move toward the retrieval process, but it is statistics about the content rather than content itself; this grants more efficient use of network resources and wider scope of application. As in distributed retrieval, some processing is distributed along with the data, but it is indexing rather than retrieval; this reduces the costs of content provision while promoting the simplicity, effectiveness, and responsiveness of retrieval. Overall, we argue that the approach retains the good properties of centralized retrieval without renouncing to cost‐effective, large‐scale resource pooling. We discuss the requirements associated with the approach and identify two strategies to deploy it on top of the OAI infrastructure. In particular, we define a minimal extension of the OAI protocol which supports the coordinated harvesting of full‐content indices and descriptive metadata for content resources. Finally, we report on the implementation of a proof‐of‐concept prototype service for multimodel content‐based retrieval of distributed file collections.
mod_oai:用于元数据收集的 Apache 模块
DOI: 10.1007/11551362_58
发表时间: 2005
期刊: Neurodegeneration : a journal for neurodegenerative disorders, neuroprotection, and neuroregeneration
影响因子: --
作者:
Michael L. Nelson;H. Sompel;Xiaoming Liu;Terry L. Harrison;Nathan McFarland
通讯作者: Nathan McFarland
修剪长文档以进行分布式信息检索
DOI: 10.1145/584792.584847
发表时间: 2002
期刊: Lit. Linguistic Comput.
影响因子: --
作者:
Jie Lu;Jamie Callan
通讯作者: Jamie Callan
DP9:网络爬虫的OAI网关服务
DOI: 10.1145/544220.544284
发表时间: 2002
期刊: D Lib Mag.
影响因子: --
作者:
Xiaoming Liu;K. Maly;M. Zubair;Michael L. Nelson
通讯作者: Michael L. Nelson
搜索中间件和简单数字图书馆互操作协议
DOI: 10.1045/march2000-paepcke
发表时间: 2000
期刊: D Lib Mag.
影响因子: --
作者:
A. Paepcke;Robert K. Brandriff;G. Janée;R. Larson;Bertram Ludäscher;S. Melnik;S. Raghavan
通讯作者: S. Raghavan
DOI: 10.1045/december2004-vandesompel
发表时间: 2004
期刊: D Lib Mag.
影响因子: --
作者:
H. Sompel;Michael L. Nelson;C. Lagoze;Simeon Warner
通讯作者: Simeon Warner