Sentence-BERT Distinguishes Good and Bad Essays in Cross-prompt Automated Essay Scoring

Sentence-BERT Distinguishes Good and Bad Essays in Cross-prompt Automated Essay Scoring
复制标题

Sentence-BERT 在交叉提示自动作文评分中区分好作文和坏作文

DOI:
10.1109/icdmw58026.2022.00045
复制
发表时间:
2022
期刊:
Proceedings of 2022 IEEE International Conference on Data Mining Workshops (ICDMW)
影响因子:
--
通讯作者:
Masada Tomonari
Masada Tomonari
中科院分区:
--
文献类型:
--
作者:
Sasaki Toru;Masada Tomonari

文献摘要

相似文献

自动作文评分(AES)指的是一套使用机器学习模型自动为学生撰写的作文评分的过程。现有的AES模型大多是即时培训的--特别是有监督的学习,这要求系统供应商在模型培训时能够访问论文提示。然而,高风险测试的作文提示通常应该在考试日期之前保密,这要求模型可以迅速交叉训练,并且已经掌握了预评分的作文数据。从诸如语句-BERT(Sbert)等预先训练的语言模型获得的文档嵌入主要被期望表示文本的语义内容。我们假设SBERT嵌入也包含与评估相关的元素,这些元素可以通过主成分分析(PCA)和归一化折扣累积收益(NDCG)度量增强的文档嵌入分解来提取。然后,在源作文的整个嵌入空间中识别出的评价元素被交叉迅速地转移到目标作文中,该目标作文写在不同的提示下,用于划分高分/低分组的二进制聚类任务。这一结果表明,非精调的SBERT已经包含了区分好与差文章的评价因素。
Automated Essay Scoring (AES) refers to a set of processes that automatically assigns grades to student-written essays with machine learning models. Existing AES models are mostly trained prompt-specifically with supervised learning, which requires the essay prompt to be accessible to the system vendor at the time of model training. However, essay prompts for high-stakes testing should usually be kept confidential before the test date, which demands the model to be cross-promptly trainable with pre-scored essay data already in hands. Document embeddings obtained from pretrained language models such as Sentence-BERT (sbert) are primarily expected to represent the semantic content of the text. We hypothesize SBERT embeddings also contain assessment-relevant elements that are extractable by document embedding decomposition through Principal Component Analysis (PCA) enhanced with Normalized Discounted Cumulative Gain (nDCG) measurement. The identified evaluative elements in the entire embedding space of the source essays are then cross-promptly transferred to the target essays written on different prompts for binary clustering task of dividing high/low-scored groups. The result implies non-finetuned SBERT already contains evaluative elements to distinguish good and bad essays.