The ACODEA framework: Developing segmentation and classification schemes for fully automatic analysis of online discussions

The ACODEA framework: Developing segmentation and classification schemes for fully automatic analysis of online discussions
复制标题

DOI:
10.1007/s11412-012-9147-y
复制
发表时间:
2012-06-01
影响因子:
4.3
通讯作者:
Fischer, Frank
Fischer, Frank
中科院分区:
教育学1区
文献类型:
--
作者:
Mu, Jin;Stegmann, Karsten;Fischer, Frank

文献摘要

被引文献

相似文献

在线讨论研究经常面临着分析大型语料库的问题。自然语言处理(NLP)技术可以允许自动化这种分析。然而,最先进的机器学习和文本挖掘方法产生的模型不能在与不同主题相关的语料库之间很好地转换。此外,分割是一个必要的步骤,但通常情况下,训练模型对训练模型时使用的分割细节非常敏感。因此,在先前发表的关于CSCL背景下的文本分类的研究中,数据是手工分割的。我们讨论如何克服这些挑战。我们提出了一个框架,开发自动分割和上下文无关的编码,建立在此分割优化的编码方案。其核心思想是在对原始数据进行切分和分类之前,利用词性标注和命名实体识别技术提取每个单词的语义和句法特征。我们的研究结果表明,在微观论证维度上的编码可以完全自动化。最后,我们讨论了如何完全自动化的分析可以使上下文敏感的支持协作学习。
Research related to online discussions frequently faces the problem of analyzing huge corpora. Natural Language Processing (NLP) technologies may allow automating this analysis. However, the state-of-the-art in machine learning and text mining approaches yields models that do not transfer well between corpora related to different topics. Also, segmenting is a necessary step, but frequently, trained models are very sensitive to the particulars of the segmentation that was used when the model was trained. Therefore, in prior published research on text classification in a CSCL context, the data was segmented by hand. We discuss work towards overcoming these challenges. We present a framework for developing coding schemes optimized for automatic segmentation and context-independent coding that builds on this segmentation. The key idea is to extract the semantic and syntactic features of each single word by using the techniques of part-of-speech tagging and named-entity recognition before the raw data can be segmented and classified. Our results show that the coding on the micro-argumentation dimension can be fully automated. Finally, we discuss how fully automated analysis can enable context-sensitive support for collaborative learning.