Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models

Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models
复制标题

DOI:
10.1145/3491101.3519665
复制
发表时间:
2022-04
期刊:
CHI Conference on Human Factors in Computing Systems Extended Abstracts
影响因子:
--
通讯作者:
Priyan Vaithilingam;Tianyi Zhang;Elena L. Glassman
Priyan Vaithilingam;Tianyi Zhang;Elena L. Glassman
中科院分区:
其他
文献类型:
--
作者:
Priyan Vaithilingam;Tianyi Zhang;Elena L. Glassman

文献摘要

被引文献

相似文献

大型语言模型(LLM)的最新进展使得在通用编程语言(例如Python)中的实际编程任务成为可能。编程工作流程发现,尽管Copilot并不一定会改善任务完成时间或成功率,但大多数参与者宁愿在日常编程任务中使用Copilot,因为Copilot经常提供了一个有用的起点并节省了在线搜索的努力。在理解,编辑和调试代码片段时,Copilot生成的代码片段极大地阻碍了他们的任务解决有效性。根据我们的观察结果和参与者的反馈,提高了改善副标士设计的有希望的方向。
Recent advances in Large Language Models (LLM) have made automatic code generation possible for real-world programming tasks in general-purpose programming languages such as Python. However, there are few human studies on the usability of these tools and how they fit the programming workflow. In this work, we conducted a within-subjects user study with 24 participants to understand how programmers use and perceive Copilot, a LLM-based code generation tool. We found that, while Copilot did not necessarily improve the task completion time or success rate, most participants preferred to use Copilot in daily programming tasks, since Copilot often provided a useful starting point and saved the effort of searching online. However, participants did face difficulties in understanding, editing, and debugging code snippets generated by Copilot, which significantly hindered their task-solving effectiveness. Finally, we highlighted several promising directions for improving the design of Copilot based on our observations and participants’ feedback.