Automatic coding of students' writing via Contrastive Representation Learning in the Wasserstein space

Automatic coding of students' writing via Contrastive Representation Learning in the Wasserstein space
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Ruijie Jiang;J. Gouvea;David Hammer;S. Aeron
Ruijie Jiang;J. Gouvea;David Hammer;S. Aeron
中科院分区:
其他
文献类型:
--
作者:
Ruijie Jiang;J. Gouvea;David Hammer;S. Aeron

文献摘要

被引文献

相似文献

对言语数据的定性分析在学习科学中具有核心重要性。然而,这是劳动密集型和耗时的,这限制了研究人员可以包括在研究中的数据量。这项工作是朝着建立统计机器学习(ML)方法迈出的一步,该方法用于实现对学生写作的定性分析的自动化支持,特别是在介绍生物学的实验室报告中,用于论证和推理的复杂性。我们从本科生物学课程的一组实验报告开始,通过四个级别的方案进行评分,该方案考虑了论点结构的复杂性,证据的范围以及结论的谨慎性和细微差别。使用这组标记数据,我们表明,一个流行的自然语言建模处理管道,即单词的矢量表示,也就是单词嵌入,然后是长短期记忆(LSTM)模型,用于捕获语言生成作为状态空间模型,能够定量捕获得分,具有高二次加权Kappa(QWK)预测得分,通过一种新的对比学习设置进行训练。我们表明,ML算法接近人的分析的评分员间的可靠性。最终,我们得出结论,用于自然语言处理(NLP)的机器学习(ML)有望帮助学习科学研究人员在比目前更大的规模上进行定性研究。
Qualitative analysis of verbal data is of central importance in the learning sciences. It is labor-intensive and time-consuming, however, which limits the amount of data researchers can include in studies. This work is a step towards building a statistical machine learning (ML) method for achieving an automated support for qualitative analyses of students' writing, here specifically in score laboratory reports in introductory biology for sophistication of argumentation and reasoning. We start with a set of lab reports from an undergraduate biology course, scored by a four-level scheme that considers the complexity of argument structure, the scope of evidence, and the care and nuance of conclusions. Using this set of labeled data, we show that a popular natural language modeling processing pipeline, namely vector representation of words, a.k.a word embeddings, followed by Long Short Term Memory (LSTM) model for capturing language generation as a state-space model, is able to quantitatively capture the scoring, with a high Quadratic Weighted Kappa (QWK) prediction score, when trained in via a novel contrastive learning set-up. We show that the ML algorithm approached the inter-rater reliability of human analysis. Ultimately, we conclude, that machine learning (ML) for natural language processing (NLP) holds promise for assisting learning sciences researchers in conducting qualitative studies at much larger scales than is currently possible.