A Duplicate Web Entity Identification Approach Based on Iterative Training

A Duplicate Web Entity Identification Approach Based on Iterative Training
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
Journal of Frontiers of Computer Science and Technology
影响因子:
--
通讯作者:
Liubao Wei
Liubao Wei
中科院分区:
其他
文献类型:
--
作者:
Liubao Wei

文献摘要

被引文献

相似文献

大量可以在线访问的Web数据源为用户获取所需信息提供了便利,作为Web数据集成的必要步骤,需要从Web数据源中准确地识别出具有不同表现形式的重复Web实体,据我们所知,以往的工作主要集中在两个数据源之间,大量的Web数据源使得这些方法不切实际。为此,提出了一种有效的基于迭代训练的重复Web实体识别方法,该方法可以利用较小的训练集应用于多个Web数据源,在图书领域和计算机领域的大量实验验证了该方法的有效性。
A large number of Web data sources that can be accessed online make users convenient to obtain their desired information.As the necessary step in Web data integration,the duplicate Web entities with various presentations should be identified accurately from Web data sources.To the best of our knowledge,previous works focus on this issue only between two data sources.The large quantity of Web data sources make these approaches unpractical. To this end,an effective iterative-training-based approach is proposed to address this issue of duplicate Web entity identification,which can be applied to multiple Web data sources using a small training set.The extensive experiments on book domain and computer domain validate the effectiveness of the proposed approach.