Search engine coverage of the OAI-PMH corpus

Search engine coverage of the OAI-PMH corpus
复制标题

OAI-PMH 语料库的搜索引擎覆盖率

DOI:
10.1109/mic.2006.41
复制
发表时间:
2006
影响因子:
3.2
通讯作者:
M. Zubair
M. Zubair
中科院分区:
计算机科学4区
文献类型:
--
作者:
F. McCown;Xiaoming Liu;Michael L. Nelson;M. Zubair

文献摘要

被引文献

相似文献

在索引了大部分“表层”网络之后,搜索引擎现在正在使用各种方法来索引“深层”网络。与此同时,机构储存库和数字图书馆正在采用元数据采集开放档案倡议协议(OAI-PMH)来公开其馆藏。作者从OAI-PMH存储库中获得了近1000万条记录。从这些记录中,他们提取了330万个唯一的资源URL,然后对这个集合中的样本进行搜索,以确定三大搜索引擎已经索引了多少OAI-PMH语料库。
Having indexed much of the "surface" Web, search engines are now using various approaches to index the "deep" Web. At the same time, institutional repositories and digital libraries are adopting the open archives initiative protocol for metadata harvesting (OAI-PMH) to expose their holdings. The authors harvested nearly 10 million records from OAI-PMH repositories. From these records, they extracted 3.3 million unique resource URLs and then conducted searches on samples from this collection to determine how much of the OAI-PMH corpus the three major search engines have indexed.