Testbed for information extraction from deep web

Testbed for information extraction from deep web
复制标题

DOI:
10.1145/1013367.1013468
复制
发表时间:
2004-05
期刊:
Synthesis
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

可搜索数据库生成的搜索结果是动态提供的,并且比Web上的静态文档大得多。这些结果页面被称为Deep Web。我们需要提取结果页面中的目标数据,以便将它们集成到不同的可搜索数据库中。我们提出了一个从搜索结果中提取信息的测试平台。我们从114,540页搜索表单中随机选择了100个数据库。因此,这些数据库的种类繁多。我们选择了51个在结果页面中包含URL的数据库,并手动识别要提取的目标信息。我们还提出了评估措施,比较提取方法和方法扩展的目标数据。
Search results generated by searchable databases are served dynamically and far larger than the static documents on the Web. These results pages have been referred to as the Deep Web. We need to extract the target data in results pages to integrate them on different searchable databases. We propose a test bed for information extraction from search results. We chose 100 databases randomly from 114,540 pages with search forms. Therefore, these databases have a good variety. We selected 51 databases which include URLs in a results pageand manually identify target information to be extracted. We also suggest evaluation measures for comparing extraction methods and methods for extending the target data.