A Quantification of Students Coding Style Utilizing HMMBased Coding Models for In-Class Source Code Plagiarism Detection

A Quantification of Students Coding Style Utilizing HMMBased Coding Models for In-Class Source Code Plagiarism Detection
复制标题

利用基于 HMM 的编码模型进行课堂源代码抄袭检测对学生编码风格的量化

DOI:
10.1109/icicic.2008.614
复制
发表时间:
2008
期刊:
2008 3rd International Conference on Innovative Computing Information and Control
影响因子:
--
通讯作者:
H. Murao
H. Murao
中科院分区:
--
文献类型:
--
作者:
Asako Ohno;H. Murao

文献摘要

被引文献

相似文献

测量在编程课中产生的源代码(以下称为“课内”源代码)之间的相似性以用于分级或检测剽窃是一项费力的任务。类内源代码的相似性度量需要一种特殊的方法,因为:(1)它们通常太短,无法提取足够的算法特征;(2)由于它们是出于相同的目的而制作的,因此它们自然具有很强的算法相似性,并且很难区分其中的剽窃和巧合相似性。本文的贡献是基于学生的编码风格而不是算法特征来量化特征。我们用一个随机模型来近似学生的编码风格,这是源代码的表面特征,称为基于隐马尔可夫模型的编码模型,并将其用于作者的认证信息。
Measuring similarity among source codes produced in programming class, hereinafter called 'in-class' source codes, for grading or detecting plagiarisms is a laborious task. A special similarity measuring method for in-class source codes is needed because: (1) they are often too short to extract enough algorithmic features, and (2) they naturally have strong algorithmic similarity since they are made for the same purpose, and it is difficult to distinguish plagiarism and coincidental similarity in them. The contribution of this paper is to quantify the features based on students' coding style instead of algorithmic features. We approximate a student's coding style which is superficial feature of a source code by a stochastic model, called coding model based on Hidden Markov Model and use it for authentification information of an author.