XML version detection
XML version detection
复制标题
XML版本检测
DOI:
10.1145/1284420.1284441
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
C. Zaniolo
中科院分区:
文献类型:
--
作者:
Deise de Brum Saccol;Nina Edelweiss;R. Galante;C. Zaniolo
The problem of version detection is critical in many important application scenarios, including software clone identification, Web page ranking, plagiarism detection, and peer-to-peer searching. A natural and commonly used approach to version detection relies on analyzing the similarity between files. Most of the techniques proposed so far rely on the use of hard thresholds for similarity measures. However, defining a threshold value is problematic for several reasons: in particular (i) the threshold value is not the same when considering different similarity functions, and (ii) it is not semantically meaningful for the user. To overcome this problem, our work proposes a version detection mechanism for XML documents based on Naïve Bayesian classifiers. Thus, our approach turns the detection problem into a classification problem. In this paper, we present the results of various experiments on synthetic data that show that our approach produces very good results, both in terms of recall and precision measures.