Multilingual Plagiarism Detection

Multilingual Plagiarism Detection
复制标题

多语言抄袭检测

DOI:
--
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
L. Agarwal
L. Agarwal
中科院分区:
--
文献类型:
--
作者:
Menna Mostafa;L. Agarwal

文献摘要

被引文献

相似文献

跨语言剽窃检测最近引起了人们的关注,因为在许多领域,如教育,新闻,科学研究,文学,电影剧本等,作者会将L1语言的文章翻译成L2语言,然后要么发表/提交它,要么改变一些句子以适应他/她的动机。因此,需要一种鲁棒的跨语言剽窃检测方法。现有的跨语言剽窃检测工作大多使用机器翻译将L1中的可疑文档翻译成L2,然后在L2中搜索相似的文档。然而,我们认为这种方法有以下局限性:(1)机器翻译不能捕捉不同领域之间不同的写作风格,(2)在线机器翻译允许匿名用户建议更好的翻译,其遭受篡改和不正确的建议,(3)识别不同类型的剽窃的能力有限,例如,两篇描述事故的文章可能被标记为剽窃,尽管它们来自不同的来源。因此,我们提出了一种方法,试图通过使用机器学习和众包技术来弥补上述三个限制。
Cross lingual plagiarism detection has recently caught attention due to copy-right violations occurring in many fields such as education, journalism, scientific research, literature, screenplays, etc, where an author would translate an article in language L1 into language L2 and then either publish/submit it or change some of the sentences to suit his/her motivations. Therefore, the need for a robust method for cross lingual plagiarism detection arises. Most of the existing work on cross lingual plagiarism detection uses machine translation to translate the suspect document in L1 into L2 and then search for similar documents in L2. However, we argue that this approach suffers from the following limitations: (1) machine translation does not capture different writing styles that differ from field to another, (2) online machine translation that allows anonymous users to suggest better translations which suffers from tampering and incorrect suggestions, (3) the limited ability to identify different types of plagiarisms, for example two articles describing an accident might be labeled as plagiarized although they originated from different sources. Therefore, we propose an approach that will attempt to remedy the above three limitations by using machine learning and crowd sourcing techniques.