Testbed for information extraction from deep web
Testbed for information extraction from deep web
复制标题
DOI:
10.1145/1013367.1013468
复制
发表时间:
2004-05
期刊:
影响因子:
--
通讯作者:
中科院分区:
文献类型:
--
作者:
Search results generated by searchable databases are served dynamically and far larger than the static documents on the Web. These results pages have been referred to as the Deep Web. We need to extract the target data in results pages to integrate them on different searchable databases. We propose a test bed for information extraction from search results. We chose 100 databases randomly from 114,540 pages with search forms. Therefore, these databases have a good variety. We selected 51 databases which include URLs in a results pageand manually identify target information to be extracted. We also suggest evaluation measures for comparing extraction methods and methods for extending the target data.