ChatGPT and Software Testing Education: Promises & Perils

ChatGPT and Software Testing Education: Promises & Perils
复制标题

DOI:
10.1109/icstw58534.2023.00078
复制
发表时间:
2023-02
期刊:
2023 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW)
影响因子:
--
通讯作者:
Sajed Jalil;Suzzana Rafi;Thomas D. Latoza;Kevin Moran;Wing Lam
Sajed Jalil;Suzzana Rafi;Thomas D. Latoza;Kevin Moran;Wing Lam
中科院分区:
其他
文献类型:
--
作者:
Sajed Jalil;Suzzana Rafi;Thomas D. Latoza;Kevin Moran;Wing Lam

文献摘要

被引文献

相似文献

在过去的十年中,代码的预测语言建模已被证明是为开发人员提供新形式的自动化的宝贵工具。最近,我们看到了基于神经变形金刚体系结构的通用“大语言模型”的广告,这些模型已经在人类书面文本的大量数据集中进行了培训,其中包括代码和自然语言。但是,尽管这种模型具有表现出的代表性,但与之相互作用在历史上一直限于特定的任务设置,从而限制了它们的一般适用性。最近,通过引入chatgpt的引入,这是一个由OpenAI创建的语言模型并受过培训,可以作为对话代理人运行,使其能够回答问题并回答最终用户的各种命令。引入模型,介绍模型,例如Chatgpt,已经激发了教育工作者的热情讨论,从担心学生可以使用这些AI工具来规避学习,到对他们可能会解锁的新型学习机会的兴奋。但是,鉴于这些工具的新生性质,我们目前缺乏与它们在不同的教育环境中的表现以及他们可能对传统教学形式所构成的潜在希望(或危险)相关的基本知识。因此,在本文中,我们研究了在受欢迎的软件测试课程中回答常见问题的任务时,Chatgpt的性能表现良好。我们发现,鉴于其当前功能,Chatppt能够回答我们研究的问题的77.5%,并且在这些问题中,它能够在55.6%的案例中提供正确或部分正确的答案,并提供正确或部分正确的解释。在53.0%的情况下的答案,并且在共同的问题上下文中促使工具的答案导致正确的答案和解释的速度较高。基于这些发现,我们讨论了与学生和讲师使用CHATGPT有关的潜在承诺和危险。
Over the past decade, predictive language modeling for code has proven to be a valuable tool for enabling new forms of automation for developers. More recently, we have seen the ad-vent of general purpose "large language models", based on neural transformer architectures, that have been trained on massive datasets of human written text, which includes code and natural language. However, despite the demonstrated representational power of such models, interacting with them has historically been constrained to specific task settings, limiting their general applicability. Many of these limitations were recently overcome with the introduction of ChatGPT, a language model created by OpenAI and trained to operate as a conversational agent, enabling it to answer questions and respond to a wide variety of commands from end users.The introduction of models, such as ChatGPT, has already spurred fervent discussion from educators, ranging from fear that students could use these AI tools to circumvent learning, to excitement about the new types of learning opportunities that they might unlock. However, given the nascent nature of these tools, we currently lack fundamental knowledge related to how well they perform in different educational settings, and the potential promise (or danger) that they might pose to traditional forms of instruction. As such, in this paper, we examine how well ChatGPT performs when tasked with answering common questions in a popular software testing curriculum. We found that given its current capabilities, ChatGPT is able to respond to 77.5% of the questions we examined and that, of these questions, it is able to provide correct or partially correct answers in 55.6% of cases, provide correct or partially correct explanations of answers in 53.0% of cases, and that prompting the tool in a shared question context leads to a marginally higher rate of correct answers and explanations. Based on these findings, we discuss the potential promises and perils related to the use of ChatGPT by students and instructors.