Am I Wrong, or Is the Autograder Wrong? Effects of AI Grading Mistakes on Learning

Am I Wrong, or Is the Autograder Wrong? Effects of AI Grading Mistakes on Learning
复制标题

DOI:
10.1145/3568813.3600124
复制
发表时间:
2023-08
期刊:
Proceedings of the 2023 ACM Conference on International Computing Education Research - Volume 1
影响因子:
--
通讯作者:
T. Li;Silas Hsu;Max Fowler;Zhilin Zhang;C. Zilles;Karrie Karahalios
T. Li;Silas Hsu;Max Fowler;Zhilin Zhang;C. Zilles;Karrie Karahalios
中科院分区:
其他
文献类型:
--
作者:
T. Li;Silas Hsu;Max Fowler;Zhilin Zhang;C. Zilles;Karrie Karahalios

文献摘要

相似文献

人工智能评分和反馈中的错误通常有一系列难以解决的原因,从本质上讲,很难完全避免。由于不准确的反馈可能会损害学习,因此需要设计和工作流程来减轻这些损害。为了更好地了解错误的人工智能反馈影响学生学习的机制,我们进行了调查和访谈,记录了学生与短答式人工智能自动评分器就用简单英语解释代码阅读问题的互动情况。使用因果模型,我们推断了标记为正确的错误答案(假阳性,FP)和标记为错误的正确答案(假阴性,FN)的学习影响。我们进一步探讨了学习影响的解释,包括影响参与者参与反馈的错误和对其答案正确性的评估,以及参与者在课堂上的先前表现。FPS在很大程度上损害了学习,这是因为参与者未能发现错误。这是因为参与者在被标记为正确后没有注意到反馈,而且一旦被标记为正确,就明显不愿承认自己的答案是错误的。另一方面,神经网络只损害了调查参与者的学习,这表明受访者更大的行为和认知投入保护了他们免受学习损害。基于这些发现,我们提出了一些方法来帮助学习者发现FP,并鼓励对FN进行更深入的反思,以减轻人工智能错误对学习的伤害。
Errors in AI grading and feedback often have an intractable set of causes and are, by their nature, difficult to completely avoid. Since inaccurate feedback potentially harms learning, there is a need for designs and workflows that mitigate these harms. To better understand the mechanisms by which erroneous AI feedback impacts students’ learning, we conducted surveys and interviews that recorded students’ interactions with a short-answer AI autograder for “Explain in Plain English” code reading problems. Using causal modeling, we inferred the learning impacts of wrong answers marked as right (false positives, FPs) and right answers marked as wrong (false negatives, FNs). We further explored explanations for the learning impacts, including errors influencing participants’ engagement with feedback and assessments of their answers’ correctness, and participants’ prior performance in the class. FPs harmed learning in large part due to participants’ failures to detect the errors. This was due to participants not paying attention to the feedback after being marked as right, and an apparent bias against admitting one’s answer was wrong once marked right. On the other hand, FNs harmed learning only for survey participants, suggesting that interviewees’ greater behavioral and cognitive engagement protected them from learning harms. Based on these findings, we propose ways to help learners detect FPs and encourage deeper reflection on FNs to mitigate the learning harms of AI errors.