Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code Summarization

Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code Summarization
复制标题

DOI:
10.1145/3643916.3644434
复制
发表时间:
2024-02
期刊:
2024 IEEE/ACM 32nd International Conference on Program Comprehension (ICPC)
影响因子:
--
通讯作者:
Jiliang Li;Yifan Zhang;Z. Karas;Collin McMillan;Kevin Leach;Yu Huang
Jiliang Li;Yifan Zhang;Z. Karas;Collin McMillan;Kevin Leach;Yu Huang
中科院分区:
其他
文献类型:
--
作者:
Jiliang Li;Yifan Zhang;Z. Karas;Collin McMillan;Kevin Leach;Yu Huang

文献摘要

相似文献

最近的语言模型已经证明能够熟练地总结源代码。然而,与机器学习的许多其他领域一样,代码的语言模型缺乏足够的非正式解释性,我们缺乏对模型从代码中学习什么以及如何学习的公式化或直观的理解。当语言模型学习生成更高质量的代码摘要时,如果它们还一致地认为与人类程序员识别的代码部分相同重要,则可以部分地提供语言模型的可解释性。在这篇文章中,我们报告了我们从人类理解的角度对代码摘要中语言模型的可解释性进行研究的负面结果。在代码摘要任务中,我们使用注视次数和持续时间等眼动跟踪指标来衡量人类对代码的关注度。为了接近语言模型的焦点,我们使用了一种最先进的模型-不可知的、黑箱的、基于扰动的方法Shap(Shapley Additive Expartions)来识别哪些代码标记影响摘要的生成。使用这些设置,我们发现语言模型的关注度和人类程序员的注意力之间没有统计上的显著关系。此外,在这种设置中,模型和人工焦点之间的对齐似乎并不决定LLM生成的摘要的质量。我们的研究突出了无法将人类的焦点与基于Shap的模型焦点测量相一致。这一结果要求进一步研究用于代码摘要和软件工程任务的可解释语言模型的多个开放问题,包括代码语言模型的训练机制,人类和模型对代码的关注是否一致,人类的关注是否可以促进语言模型的发展,以及其他哪些模型重点措施适合于提高可解释性。CCS概念·计算方法$\rright$人工智能。
Recent language models have demonstrated proficiency in summarizing source code. However, as in many other domains of machine learning, language models of code lack sufficient explainability informally, we lack a formulaic or intuitive understanding of what and how models learn from code. Explainability of language models can be partially provided if, as the models learn to produce higher-quality code summaries, they also align in deeming the same code parts important as those identified by human programmers. In this paper, we report negative results from our investigation of explainability of language models in code summarization through the lens of human comprehension. We measure human focus on code using eye-tracking metrics such as fixation counts and duration in code summarization tasks. To approximate language model focus, we employ a state-of-the-art model-agnostic, black-box, perturbation-based approach, SHAP (SHapley Additive exPlanations), to identify which code tokens influence that generation of summaries. Using these settings, we find no statistically significant relationship between language models’ focus and human programmers’ attention. Furthermore, alignment between model and human foci in this setting does not seem to dictate the quality of the LLM-generated summaries. Our study highlights an inability to align human focus with SHAP-based model focus measures. This result calls for future investigation of multiple open questions for explainable language models for code summarization and software engineering tasks in general, including the training mechanisms of language models for code, whether there is an alignment between human and model attention on code, whether human attention can improve the development of language models, and what other model focus measures are appropriate for improving explainability. CCS CONCEPTS • Computing methodologies $\rightarrow$ Artificial intelligence.