Classification of Programming Problems based on Topic Modeling

Classification of Programming Problems based on Topic Modeling
复制标题

DOI:
10.1145/3323771.3323795
复制
发表时间:
2019-03
期刊:
Proceedings of the 2019 7th International Conference on Information and Education Technology - ICIET 2019
影响因子:
--
通讯作者:
Chowdhury Md Intisar;Y. Watanobe;Manoj Poudel;S. Bhalla
Chowdhury Md Intisar;Y. Watanobe;Manoj Poudel;S. Bhalla
中科院分区:
其他
文献类型:
--
作者:
Chowdhury Md Intisar;Y. Watanobe;Manoj Poudel;S. Bhalla

文献摘要

被引文献

相似文献

编程技能是当今时代最重要和要求最高的技能之一。为了使学习者和程序员能够练习编程并获得解决问题的技能,存在许多在线评判(OJ)系统。这些OJ系统中的大多数必须由学生和学习者单独操作。这些学生和新手程序员有时会互相竞争,或者在离线模式下自己解决编程问题。但是,大多数OJ系统的问题只是简单地安排成卷和各种比赛事件。这个安排制度并没有清楚显示问题的困难和类别。因此,在本文中,我们已经研究了可靠的关键字和功能,可以将这些OJ系统的编程问题到各自的类型和技能的提取技术。我们利用两种流行的主题建模算法,潜在狄利克雷分配(LDA)和非负矩阵分解(NMF)来提取相关特征。之后,六个分类器被训练这些主题建模特征和朴素TF-IDF特征。从我们的研究中,我们发现主题建模特征在维度上相对较小,但在高维朴素TF-IDF特征上训练时匹配性能。我们的主要目标是理解编程问题语句的文本数据的准确性和维度之间的精确权衡。通过实验,我们获得了在线判断编程问题的重要标记、提示和分类。
Programming skill is one of the most important and demanding skill in the current generation. In order to enable learners and programmers to practice programming and gain problem-solving skills, many Online Judge (OJ) systems exist. Most of these OJ systems have to be operated solely by students and learners. These students and novice programmers sometimes compete against each other or solve the programming problems by themselves in offline mode. But, most OJ systems have their problems arranged simply into volumes and various contests events. This arrangement system does not have any clear indication of the difficulties and categories of problems. Thus, in this paper, we have studied reliable techniques on the extraction of keywords and features which can categorize these OJ system's programming problems into their respective types and skills. We have leveraged two popular topic modeling algorithms, Latent Dirichlet Allocation (LDA) and Non-negative Matrix Factorization (NMF) to extract relevant features. Afterward, six classifiers were trained on these topic modeling features and Naive TF-IDF features. From our studies, we discovered that topic modeling features were relatively smaller in dimensionality, yet matched the performance when trained on high dimensional naive TF-IDF features. Our main goal was to understand the precise trade-off between accuracy and dimensionality of the textual data of programming problem statements. This experiment has enabled us to obtain important tags, hint, and classification of Online Judge programming problems.