A Plagiarism Detection System for Arabic Text-Based Documents

A Plagiarism Detection System for Arabic Text-Based Documents
复制标题

阿拉伯文本文档的抄袭检测系统

DOI:
10.1007/978-3-642-30428-6_12
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Ashraf Elnagar
Ashraf Elnagar
中科院分区:
--
文献类型:
--
作者:
A. Jadalla;Ashraf Elnagar

文献摘要

被引文献

相似文献

本文提出了一种新的剽窃检测系统,阿拉伯语文本为基础的文件,Iqtebas 1.0。这是一个主要的工作致力于剽窃阿拉伯文为基础的文件。阿拉伯语是一种丰富的形态语言,是世界上最常用的语言之一,也是互联网上最常用的语言。给定一个文档和一组可疑文件,我们的目标是计算被检查文档的原始值。通过计算文本中的每个句子与可疑文件中最近的句子之间的距离来计算文本的原创性值。建议的系统结构是基于一个搜索引擎,以减少成对相似性的成本。在索引过程中,我们使用了精选n-gram指纹算法来减少索引的大小。每个句子的指纹是由哈希码表示的n-gram。筛选算法计算每个句子的指纹。因此,搜索时间得到改善,并且检测过程是准确和鲁棒的。实验结果表明,Iqtebas 1.0的查全率和查准率分别达到了94%和99%;并与著名的剽窃检测系统SafeAssign进行了比较,证实了Iqtebas的高性能。
This paper presents a novel plagiarism detection system for Arabic text-based documents, Iqtebas 1.0. This is a primary work dedicated for plagiarism of Arabic based documents. Arabic is a rich morphological language that is among the top used languages in the world and in the Internet as well. Given a document and a set of suspected files, our goal is to compute the originality value of the examined document. The originality value of a text is computed by computing the distance between each sentence in the text and the closest sentence in the suspected files, if exists. The proposed system structure is based on a search engine in order to reduce the cost of pairwise similarity. For the indexing process, we use the winnowing n-gram fingerprinting algorithm to reduce the index size. The fingerprints of each sentence are its n-grams that are represented by hash codes. The winnowing algorithm computes fingerprints for each sentence. As a result, the search time is improved and the detection process is accurate and robust. The experimental results showed superb performance of Iqtebas 1.0 as it achieved a recall value of 94% and a precision of 99%.Moreover, a comparison that is carried out between Iqtebas and the well known plagiarism detection system, SafeAssign, confirmed the high performance of Iqtebas.