More With Less: Exploring How to Use Deep Learning Effectively through Semi-supervised Learning for Automatic Bug Detection in Student Code

More With Less: Exploring How to Use Deep Learning Effectively through Semi-supervised Learning for Automatic Bug Detection in Student Code
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Yang Shi;Ye Mao;T. Barnes;Min Chi;T. Price
Yang Shi;Ye Mao;T. Barnes;Min Chi;T. Price
中科院分区:
其他
文献类型:
--
作者:
Yang Shi;Ye Mao;T. Barnes;Min Chi;T. Price

文献摘要

相似文献

自动检测学生程序代码中的错误对于实现形成性反馈以帮助学生找出错误并解决它们至关重要。深度学习模型,特别是code 2 vec和ASTNN,在大规模代码分类方面取得了巨大成功。然而,当标记数据量有限时,它们是否能有效地用于错误检测还不清楚。在这项工作中,我们研究了code 2 vec和ASTNN对经典机器学习模型的有效性,方法是将标记数据的数量从1%变化到100%。除了少数例外,这两种深度学习模型的表现优于经典模型。更有趣的是,我们的结果表明,当标记数据量较小时,code 2 vec更有效,而ASTNN在训练数据较多时更有效;对于code 2 vec和ASTNN,标记数据越多越好。为了进一步提高它们的效率,我们研究了半监督学习的潜力,它可以利用大量未标记的数据来提高它们的性能。我们的研究结果表明,半监督学习确实是贝内的,特别是对于ASTNN。
Automatically detecting bugs in student program code is critical to enable formative feedback to help students pin-point errors and resolve them. Deep learning models especially code2vec and ASTNN have shown great success for large-scale code classification. It is not clear, however, whether they can be effectively used for bug detection when the amount of labeled data is limited. In this work, we investigated the effectiveness of code2vec and ASTNN against classic machine learning models by varying the amount of labeled data from 1% up to 100%. With a few exceptions, the two deep learning models outperform the classic models. More interestingly, our results showed that when the amount of labeled data is small, code2vec is more effective, while ASTNN is more effective with more training data; for both code2vec and ASTNN, the more labeled data, the better. To further improve their effectiveness, we investigated the potential of semi-supervised learning which can leverage a large amount of unlabeled data to improve their performance. Our results showed that semi-supervised learning is indeed beneficial especially for ASTNN.