Research on Automatic Classification for Deep Web Query Interfaces

Research on Automatic Classification for Deep Web Query Interfaces
复制标题

深网查询接口自动分类研究

DOI:
10.1109/isip.2008.140
复制
发表时间:
2008
期刊:
2008 International Symposiums on Information Processing
影响因子:
--
通讯作者:
Chao Lv
Chao Lv
中科院分区:
--
文献类型:
--
作者:
Peiguang Lin;Y. Du;Xiaohua Tan;Chao Lv

文献摘要

被引文献

相似文献

近年来,Web的“深度”迅速扩大,用户要访问特定领域的Web数据库,需要浏览大量的Web站点。因此,建立一个统一的查询接口,集成一个领域的查询接口,同时访问不同的Web数据库成为一个非常重要的问题。本文首先分析了查询接口的模式特征和同一领域中的公共属性,给出了一种新的查询接口表示方法,然后提出了“形式项”和“功能项”的定义,并在此基础上提出了一种新的相似度计算算法--基于文字和语义的相似度计算(LSSC)。其次,结合LSSC和NQ算法,给出了一种Deep Web查询接口的聚类算法:LSSC-NQ。实验表明,该算法能够准确地计算出查询接口之间的相似度,并能高效、可靠、快速地对查询接口进行聚类。
In recent years, the Web is "deepened" rapidly and users have to browse quantities of Web sites to access Web databases in a specific domain. So, to build an unified query interface which integrates query interfaces of a domain to access various Web databases at the same time becomes a very important issue. In this paper, the schema characteristics of query interfaces and common attributes in a same domain are firstly analyzed, and it also gives a new representation of query interface, then the definition of "Form term" and "Function term" are proposed ,and a new similarity computing algorithm, literal and semantic based similarity computing (LSSC) is proposed, which is based on the two definitions. Secondly, a clustering algorithm for Deep Web query interfaces is given by combining LSSC and NQ algorithm: LSSC-NQ. Finally, experiments show that this algorithm can give accurate similarity computing, and cluster query interfaces efficiently, reliably and quickly.