JAG: A Crowdsourcing Framework for Joint Assessment and Peer Grading

JAG: A Crowdsourcing Framework for Joint Assessment and Peer Grading
复制标题

JAG:联合评估和同行评分的众包框架

DOI:
10.1609/aaai.v31i1.10631
复制
发表时间:
2017
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Christoph Studer
Christoph Studer
中科院分区:
--
文献类型:
--
作者:
I. Labutov;Christoph Studer

文献摘要

被引文献

相似文献

众包内容的产生和评估通常被视为两个独立的过程,在不同的时间由两个不同的人群体执行:内容创建者和内容评估者。因此,大多数众包任务都遵循这个模板:一组工作人员生成内容,另一组工作人员对其进行评估。例如,在教育环境中,内容创建者传统上是提交作业的开放式答案的学生(例如,简短的答案、电路图或公式),而内容评价员是为这些提交的内容评分的讲师。尽管同行评分在大规模在线公开课(MOOC)中取得了相当大的成功,但考试和评分过程仍然被视为两项不同的任务,通常发生在不同的时间,需要额外的评分员培训和激励费用。受教育背景下这一问题的启发,我们提出了一个通用的众包框架,它将开放回答考试(内容生成)和评估融合到一个单一的、简化的过程中,对学生来说,这个过程似乎是以显性测试的形式出现的,但每个人都扮演着隐含的评分者的角色。我们的框架提供的优势包括:创建和评估内容的共同激励机制,以及联合建模投稿和评估过程的概率模型,有助于有效估计投稿的质量和投稿者的能力。我们通过模拟和真实世界的用户研究证明了我们的框架的有效性和局限性。
Generation and evaluation of crowdsourced content is commonly treated as two separate processes, performed at different times and by two distinct groups of people: content creators and content assessors. As a result, most crowdsourcing tasks follow this template: one group of workers generates content and another group of workers evaluates it. In an educational setting, for example, content creators are traditionally students that submit open-response answers to assignments (e.g., a short answer, a circuit diagram, or a formula) and content assessors are instructors that grade these submissions. Despite the considerable success of peer-grading in massive open online courses (MOOCs), the process of test-taking and grading are still treated as two distinct tasks which typically occur at different times, and require an additional overhead of grader training and incentivization. Inspired by this problem in the context of education, we propose a general crowdsourcing framework that fuses open-response test-taking (content generation) and assessment into a single, streamlined process that appears to students in the form of an explicit test, but where everyone also acts as an implicit grader. The advantages offered by our framework include: a common incentive mechanism for both the creation and evaluation of content, and a probabilistic model that jointly models the processes of contribution and evaluation, facilitating efficient estimation of the quality of the contributions and the competency of the contributors. We demonstrate the effectiveness and limits of our framework via simulations and a real-world user study.