XML version detection

XML version detection
复制标题

XML版本检测

DOI:
10.1145/1284420.1284441
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
C. Zaniolo
C. Zaniolo
中科院分区:
--
文献类型:
--
作者:
Deise de Brum Saccol;Nina Edelweiss;R. Galante;C. Zaniolo

文献摘要

被引文献

相似文献

在许多重要的应用场景中,版本检测问题是至关重要的,包括软件克隆识别、网页排名、抄袭检测和对等搜索。一种自然且常用的版本检测方法依赖于分析文件之间的相似性。到目前为止,提出的大多数技术都依赖于使用硬阈值来进行相似性度量。然而,由于以下几个原因,定义阈值是有问题的:尤其是(I)当考虑不同的相似性函数时,阈值不相同,以及(Ii)它对用户没有语义意义。为了解决这个问题,我们的工作提出了一种基于朴素贝叶斯分类器的XML文档版本检测机制。因此,我们的方法将检测问题转化为分类问题。在本文中,我们给出了在合成数据上的各种实验结果,表明我们的方法在召回率和查准率方面都得到了很好的结果。
The problem of version detection is critical in many important application scenarios, including software clone identification, Web page ranking, plagiarism detection, and peer-to-peer searching. A natural and commonly used approach to version detection relies on analyzing the similarity between files. Most of the techniques proposed so far rely on the use of hard thresholds for similarity measures. However, defining a threshold value is problematic for several reasons: in particular (i) the threshold value is not the same when considering different similarity functions, and (ii) it is not semantically meaningful for the user. To overcome this problem, our work proposes a version detection mechanism for XML documents based on Naïve Bayesian classifiers. Thus, our approach turns the detection problem into a classification problem. In this paper, we present the results of various experiments on synthetic data that show that our approach produces very good results, both in terms of recall and precision measures.