Discovering Important Source Code Terms

Discovering Important Source Code Terms
复制标题

发现重要的源代码术语

DOI:
--
复制
发表时间:
2016
期刊:
2016 IEEE/ACM 38th International Conference on Software Engineering Companion (ICSE-C)
影响因子:
--
通讯作者:
Collin McMillan
Collin McMillan
中科院分区:
--
文献类型:
--
作者:
Paige Rodeghero;Collin McMillan

文献摘要

被引文献

相似文献

源代码中的术语在软件工程研究中变得极其重要。这些“重要”术语通常用作研究工具的输入。因此,这些工具的输出质量将取决于术语提取技术的质量。目前,还没有明确的最佳技术来预测程序理解过程中术语的重要性。在我的工作中,我对几种技术进行了文献综述。然后,我提出了一个基于机器学习算法的统一重要性预测模型。我在涉及专业程序员的实地研究以及标准的 10 倍综合研究中评估了我的模型。我发现我的模型以大约 50% 的精度和召回率预测最重要源代码术语的前四分之一,优于 tf/idf 和其他流行技术。此外,我发现,在实际的程序理解任务中,我的模型的预测可以帮助程序员相当于一组真实的最重要的术语。
Terms in source code have become extremely important in Software Engineering research. These ``important' terms are typically used as input to research tools. Therefore, the quality of the output of these tools will depend on the quality of the term extraction technique. Currently, there is no definitive best technique for predicting the importance of terms during program comprehension. In my work, I perform a literature review of several techniques. I then propose a unified importance prediction model based on a machine learning algorithm. I evaluate my model in a field study involving professional programmers, as well as a standard 10-fold synthetic study. I found that my model predicts the top quartile of most-important source code terms with approximately 50\% precision and recall, outperforming tf/idf and other popular techniques. Furthermore, I found that, during actual program comprehension tasks, the predictions from my model help programmers equivalent to a real set of most-important terms.